Budget: real ratio ledger/2.7 (Ollama page 24.63 USD)

This commit is contained in:
Kral
2026-10-03 20:11:45 +02:00
parent aef4c6e461
commit b0fd06259b
2 changed files with 7 additions and 7 deletions

View File

@@ -45,7 +45,7 @@
- A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between
sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this.
Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs.
- Budget: `runs/ledger.jsonl` (list prices, upper bound; real cost ≈ ledger / 1.47, measured 2026-10-03). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`.
- Budget: `runs/ledger.jsonl` (list prices, upper bound; real cost ≈ ledger / 2.7, measured 2026-10-03 20:15: ledger 66.97 vs Ollama monthly usage 24.63 USD (earlier ratio 1.47 was wrong)). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`.
Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work.
- Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_`
(`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`).

View File

@@ -91,10 +91,10 @@ base model (step 2) needs step 1. Details and next steps: `train/STATE.md`.
## 5. Budget
- Cycle budget 60 USD until 12 October 2026. Ledger 63.3 USD (list prices) ≈ 43 USD real (ratio 1.47,
measured 2026-10-03). About 17 USD real left.
- Guard: `BUDGET_LIMIT_USD=75` (ledger) ≈ 51 USD real.
- Tonight: about 25 DeepSeek reruns (≈ 1.5 USD real) and judge calls for H tasks in the baseline (small).
- Real usage (Ollama page, Kral, 2026-10-03 20:15): 24.63 of 60 USD. Ledger (list prices): 66.97 USD.
New ratio: real ≈ ledger / 2.7 (the old ratio 1.47 was wrong). About 35 USD real left until 12 October.
- Guard: `BUDGET_LIMIT_USD=75` (ledger) is now about 28 USD real: too strict. Proposal: raise it so that the
real cap is about 50 of 60 USD: ledger ≈ 135. Not changed yet; Kral decides.
- Planned cloud work is small: reruns (G0119, G0162, ≈ 0.15 USD real) and judge calls for H tasks in the baseline.
The baseline itself is local and free.
- The agent now limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of
the run cost.
- The agent limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of the run cost.