Budget guard: ratio 2.07 measured, limit 142; limit read from .env at each check

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 14:47:54 +02:00
parent c062df2ec5
commit ae37617a0d
2 changed files with 18 additions and 1 deletions

View File

@@ -34,9 +34,24 @@ def spent(since=None):
return round(sum(json.loads(l)["usd"] for l in open(LEDGER) if json.loads(l)["day"] >= since), 4) return round(sum(json.loads(l)["usd"] for l in open(LEDGER) if json.loads(l)["day"] >= since), 4)
def _env_budget():
"""BUDGET_LIMIT_USD and BUDGET_RESERVE_USD from .env (read at each check: a change applies at once)."""
vals = {"BUDGET_LIMIT_USD": os.environ.get("BUDGET_LIMIT_USD", "45"),
"BUDGET_RESERVE_USD": os.environ.get("BUDGET_RESERVE_USD", "0")}
try:
for line in open(os.path.join(ROOT, ".env")):
k, _, v = line.strip().partition("=")
if k in vals and v:
vals[k] = v
except OSError:
pass
return float(vals["BUDGET_LIMIT_USD"]), float(vals["BUDGET_RESERVE_USD"])
def check_budget(): def check_budget():
# BUDGET_RESERVE_USD: ledger USD kept back (second teacher test later); the guard stops at limit - reserve # BUDGET_RESERVE_USD: ledger USD kept back (second teacher test later); the guard stops at limit - reserve
limit = float(os.environ.get("BUDGET_LIMIT_USD", "45")) - float(os.environ.get("BUDGET_RESERVE_USD", "0")) limit_raw, reserve = _env_budget()
limit = limit_raw - reserve
s = spent() s = spent()
if s >= limit: if s >= limit:
raise BudgetExceeded(f"cycle spend {s} USD >= limit {limit} USD") raise BudgetExceeded(f"cycle spend {s} USD >= limit {limit} USD")

View File

@@ -142,3 +142,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation. - B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget. - 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.
- 2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = **2.07**, not 2.7. `BUDGET_LIMIT_USD` 161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept). `harness/ledger.py` now reads the limit from `.env` at each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs in `runs/traj/_aborted/`, rerun later).