Budget guard: ratio 2.07 measured, limit 142; limit read from .env at each check

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 14:47:54 +02:00
parent c062df2ec5
commit ae37617a0d
2 changed files with 18 additions and 1 deletions

View File

@@ -142,3 +142,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.
- 2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = **2.07**, not 2.7. `BUDGET_LIMIT_USD` 161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept). `harness/ledger.py` now reads the limit from `.env` at each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs in `runs/traj/_aborted/`, rerun later).