Budget guard: ratio 2.07 measured, limit 142; limit read from .env at each check
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -34,9 +34,24 @@ def spent(since=None):
|
|||||||
return round(sum(json.loads(l)["usd"] for l in open(LEDGER) if json.loads(l)["day"] >= since), 4)
|
return round(sum(json.loads(l)["usd"] for l in open(LEDGER) if json.loads(l)["day"] >= since), 4)
|
||||||
|
|
||||||
|
|
||||||
|
def _env_budget():
|
||||||
|
"""BUDGET_LIMIT_USD and BUDGET_RESERVE_USD from .env (read at each check: a change applies at once)."""
|
||||||
|
vals = {"BUDGET_LIMIT_USD": os.environ.get("BUDGET_LIMIT_USD", "45"),
|
||||||
|
"BUDGET_RESERVE_USD": os.environ.get("BUDGET_RESERVE_USD", "0")}
|
||||||
|
try:
|
||||||
|
for line in open(os.path.join(ROOT, ".env")):
|
||||||
|
k, _, v = line.strip().partition("=")
|
||||||
|
if k in vals and v:
|
||||||
|
vals[k] = v
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
return float(vals["BUDGET_LIMIT_USD"]), float(vals["BUDGET_RESERVE_USD"])
|
||||||
|
|
||||||
|
|
||||||
def check_budget():
|
def check_budget():
|
||||||
# BUDGET_RESERVE_USD: ledger USD kept back (second teacher test later); the guard stops at limit - reserve
|
# BUDGET_RESERVE_USD: ledger USD kept back (second teacher test later); the guard stops at limit - reserve
|
||||||
limit = float(os.environ.get("BUDGET_LIMIT_USD", "45")) - float(os.environ.get("BUDGET_RESERVE_USD", "0"))
|
limit_raw, reserve = _env_budget()
|
||||||
|
limit = limit_raw - reserve
|
||||||
s = spent()
|
s = spent()
|
||||||
if s >= limit:
|
if s >= limit:
|
||||||
raise BudgetExceeded(f"cycle spend {s} USD >= limit {limit} USD")
|
raise BudgetExceeded(f"cycle spend {s} USD >= limit {limit} USD")
|
||||||
|
|||||||
@@ -142,3 +142,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
|||||||
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
|
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
|
||||||
|
|
||||||
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.
|
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.
|
||||||
|
|
||||||
|
- 2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = **2.07**, not 2.7. `BUDGET_LIMIT_USD` 161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept). `harness/ledger.py` now reads the limit from `.env` at each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs in `runs/traj/_aborted/`, rerun later).
|
||||||
|
|||||||
Reference in New Issue
Block a user