diff --git a/train/STATE.md b/train/STATE.md index 1669c72..7defbd2 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -144,3 +144,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training. - 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget. - 2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = **2.07**, not 2.7. `BUDGET_LIMIT_USD` 161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept). `harness/ledger.py` now reads the limit from `.env` at each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs in `runs/traj/_aborted/`, rerun later). + +- 2026-10-05 15:55 budget correction 2: panel 34.79 at ledger 85.45. Ledger per usage is not constant: 2.04 (24.63 to 31.48, mostly generation) and 1.35 (31.48 to 34.79, mostly trajectory runs). Trajectory runs cost more real usage per ledger USD. Guard set from the panel: panel left 60 - 34.79 - 3.0 reserve = 22.2 usage x 1.35 = 30 ledger, so guard 115 ledger, `BUDGET_LIMIT_USD` 123 (reserve 8). Rule: Kral gives the panel value now and then; the limit is recomputed from it.