90 lines
7.4 KiB
Markdown
90 lines
7.4 KiB
Markdown
# Restart plan for 12 October 2026 (cloud work)
|
|
|
|
Decision (Kral + Opus, 2026-10-05): when the budget guard stops, no cloud work (generation, trajectories) before the
|
|
Ollama reset on 12 October. Until then only no-cloud work. This is the plan for the restart.
|
|
|
|
## 1. Before the start (5 minutes, no cloud call)
|
|
|
|
1. Panel value after the reset (usage USD). Then:
|
|
`python3 -m harness.restart_plan --panel <value>`
|
|
It prints the `.env` values and the state below with live numbers. Put them in `.env`:
|
|
`BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD=<printed>`, `BUDGET_RESERVE_USD=8` (the 3 usage reserve for the
|
|
second teacher test stays). Ratio for the guard: 1.2 ledger per usage (trajectory runs; generation is about 2, so the
|
|
guard is on the safe side). Give the panel value again after about 2 hours and recompute (`python3 -m harness.dashboard panel <v>`).
|
|
2. A4H up (`docker ps`), MCP answers, no stale lock: `python3 scripts_probe/lockprobe.py` must print 0 entries.
|
|
3. One controller only: `python3 -m harness.pipeline` refuses to start a second one (`runs/pipeline/controller.lock`).
|
|
Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist.
|
|
4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`).
|
|
|
|
## 1b. Step 0: eval tasks for the new kinds (before any training task)
|
|
|
|
`python3 -m harness.evalset plan-new` shows 32 slots (INTF 10, TABL 11, STRU 3, MSAG 4, exception 4; ids G0200 to G0231), then `python3 -m harness.evalset run-new`
|
|
(about 5 ledger, about 2 usage), then `python3 -m harness.evalset run-new-k` (5 K variants). Reason: the eval set (130 accepted candidates) has no INTF, TABL, STRU, MSAG
|
|
or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: `docs/eval-spotcheck-new-kinds.md`
|
|
(Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.
|
|
|
|
## 1c. Step 0b: rerun of four empirical filter runs (Kral + Opus 2026-10-06)
|
|
|
|
Four eval candidates saw or read leftover objects of other runs in the first empirical filter (`docs/foreign-objects-report.md`): **G0183, G0180, G0143, G0185** (DeepSeek, scores 80, 85, 85, 85).
|
|
After step 0 (eval generation for the new kinds), with the fixed proxy:
|
|
|
|
`python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 450000 --rerun G0183 G0180 G0143 G0185`
|
|
|
|
This writes only to `runs/emp_rerun/` (4 runs, about 1 ledger) and prints old and new filter decision per task. `empirical.json` and the eval set are **not** touched. Decision rules (step F):
|
|
easy candidate = score 95 or more and at most 20 tool calls; flag = score under 50 (a strong model fails although the reference passes); else normal. **If a decision changes for one of the four, report it to Kral and
|
|
Opus first; the eval set is changed only after their answer.** If nothing changes, say so in `docs/eval-inceleme.md` and leave the old results.
|
|
|
|
## 2. Order of work (what the controller does by itself)
|
|
|
|
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest
|
|
deficit goes first. Target (percent of accepted tasks and of accepted trajectories): CLAS 28, INTF 7, DDLS 25, FUNC 15,
|
|
PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5.
|
|
|
|
1. Trajectories for the new-type tasks that wait (INTF, TABL, STRU, MSAG, exception, PROG), then DDLS.
|
|
2. Second attempts: a task whose first attempt failed, or was accepted without a repair, gets one more attempt (at most two
|
|
attempts, at most two accepted trajectories per task). Failed DDLS and failed new-type tasks come first because they are
|
|
below target.
|
|
3. Generation (3 workers): the kind with the biggest deficit; a kind with more than 8 waiting tasks is not generated; a kind with
|
|
6 or more tries and under 20 % accepted is skipped. 20 % error-targeted slots, 30 % hard slots.
|
|
4. CLAS and FUNC generation and trajectories wait until they are at or below their target share (CLAS is far above).
|
|
5. K variants (free text) resume at 10 % of the other accepted tasks when their kinds are below target.
|
|
|
|
## 3. How much is needed (numbers of 2026-10-05 20:00, live: `restart_plan`)
|
|
|
|
CLAS has 49 accepted trajectories. At a 28 % share that is a total of 175 accepted trajectories, so about 105 more are
|
|
needed, all on other kinds: INTF 12, DDLS about 35, FUNC about 15, PROG about 17, TABL 14, STRU 3, MSAG 4, exception 4.
|
|
Tasks needed (accepted, with about 1.3 trajectory runs per accepted trajectory): TABL +9, INTF +10, STRU +3, MSAG +4,
|
|
DDLS +20, PROG +5. About 130 trajectory runs at 0.25 ledger = 33 ledger = 27 usage, plus the generation (about 6 usage).
|
|
Second attempts are part of the 130 runs. If the new reset gives 60 usage, this fits with room for the second teacher test.
|
|
|
|
## 4. Checks on the first runs of each new type (do not skip)
|
|
|
|
- The first 5 runs of INTF, TABL, STRU, MSAG, exception: records complete (the proxy must pass `sap_push_message`),
|
|
Qwen conversion works (`train/to_qwen.py` on `accepted.jsonl`: tool call round trip), no harness event.
|
|
- Acceptance per kind in the summary; a kind under 20 % after 6 tries is skipped automatically: read why (prompt, harness, or the
|
|
teacher cannot do it) before starting it again.
|
|
- DDLS: 8 of 17 runs accepted on 2026-10-05; the rejections were budget (60 calls, now 100) and empty responses (32k output limit).
|
|
If empty responses stay high, lower the output limit for CDS runs or retry once more.
|
|
|
|
## 5. Settings that stay
|
|
|
|
- Trajectory workers 2 (a harness event sets 1); stop rules: acceptance under 50 % over the last 30 runs, the same harness
|
|
error three times, the budget guard. Generation deadline 2026-10-10 18:00 has passed: set `--gen-deadline` and
|
|
`--traj-deadline` on the controller command (for example `--gen-deadline 2026-10-20T18:00 --traj-deadline 2026-10-21T23:30`).
|
|
- Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
|
|
- Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit.
|
|
|
|
## 5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks)
|
|
1. `git status` clean and pushed; `python3 -m harness.restart_plan --panel <value>` prints the numbers (panel value from Kral, expected 0 after the reset).
|
|
2. Services: A4H up, MCP answers, `python3 scripts_probe/lockprobe.py` shows no stale lock (only ATC runtime and debugger listener entries), `python3 -m harness.sweep` shows no leftover.
|
|
3. `.env`: `BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD`, `BUDGET_RESERVE_USD=8`; no `STOP` flag, no `STOPPED.txt`; no running controller (`runs/pipeline/controller.lock`).
|
|
4. Order of the first hour: step 0 eval slots for the new kinds (`evalset plan-new`, `run-new`, `run-new-k`), then the pending tasks of the kinds below target. The 9 trajectories that were
|
|
moved back (`runs/traj/summary_excluded.jsonl`: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order.
|
|
5. Settings to confirm: workers 2, `STREAM_GUARD` unset for the first runs and then 9000 for 5 DDLS/PROG runs (`docs/empty-response.md`), tool budget 100 for CDS.
|
|
6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host.
|
|
|
|
## 6. After the restart
|
|
|
|
Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the
|
|
reserve), the EPOD requests in `docs/epod-lock-leak.md` and `docs/epod-syntax-hint.md`.
|