Files
abap-llm/docs/restart-12-oktober.md

4.5 KiB

Restart plan for 12 October 2026 (cloud work)

Decision (Kral + Opus, 2026-10-05): when the budget guard stops, no cloud work (generation, trajectories) before the Ollama reset on 12 October. Until then only no-cloud work. This is the plan for the restart.

1. Before the start (5 minutes, no cloud call)

  1. Panel value after the reset (usage USD). Then: python3 -m harness.restart_plan --panel <value> It prints the .env values and the state below with live numbers. Put them in .env: BUDGET_CYCLE_START=2026-10-12, BUDGET_LIMIT_USD=<printed>, BUDGET_RESERVE_USD=8 (the 3 usage reserve for the second teacher test stays). Ratio for the guard: 1.2 ledger per usage (trajectory runs; generation is about 2, so the guard is on the safe side). Give the panel value again after about 2 hours and recompute (python3 -m harness.dashboard panel <v>).
  2. A4H up (docker ps), MCP answers, no stale lock: python3 scripts_probe/lockprobe.py must print 0 entries.
  3. One controller only: python3 -m harness.pipeline refuses to start a second one (runs/pipeline/controller.lock). Remove runs/pipeline/STOP and STOPPED.txt if they exist.
  4. Dashboard: python3 -m harness.dashboard (5-minute page, runs/dashboard/index.html).

2. Order of work (what the controller does by itself)

Only kinds below their target share are generated and run (harness/mix.py, below_target); the kind with the biggest deficit goes first. Target (percent of accepted tasks and of accepted trajectories): CLAS 28, INTF 7, DDLS 25, FUNC 15, PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5.

  1. Trajectories for the new-type tasks that wait (INTF, TABL, STRU, MSAG, exception, PROG), then DDLS.
  2. Second attempts: a task whose first attempt failed, or was accepted without a repair, gets one more attempt (at most two attempts, at most two accepted trajectories per task). Failed DDLS and failed new-type tasks come first because they are below target.
  3. Generation (3 workers): the kind with the biggest deficit; a kind with more than 8 waiting tasks is not generated; a kind with 6 or more tries and under 20 % accepted is skipped. 20 % error-targeted slots, 30 % hard slots.
  4. CLAS and FUNC generation and trajectories wait until they are at or below their target share (CLAS is far above).
  5. K variants (free text) resume at 10 % of the other accepted tasks when their kinds are below target.

3. How much is needed (numbers of 2026-10-05 20:00, live: restart_plan)

CLAS has 49 accepted trajectories. At a 28 % share that is a total of 175 accepted trajectories, so about 105 more are needed, all on other kinds: INTF 12, DDLS about 35, FUNC about 15, PROG about 17, TABL 14, STRU 3, MSAG 4, exception 4. Tasks needed (accepted, with about 1.3 trajectory runs per accepted trajectory): TABL +9, INTF +10, STRU +3, MSAG +4, DDLS +20, PROG +5. About 130 trajectory runs at 0.25 ledger = 33 ledger = 27 usage, plus the generation (about 6 usage). Second attempts are part of the 130 runs. If the new reset gives 60 usage, this fits with room for the second teacher test.

4. Checks on the first runs of each new type (do not skip)

  • The first 5 runs of INTF, TABL, STRU, MSAG, exception: records complete (the proxy must pass sap_push_message), Qwen conversion works (train/to_qwen.py on accepted.jsonl: tool call round trip), no harness event.
  • Acceptance per kind in the summary; a kind under 20 % after 6 tries is skipped automatically: read why (prompt, harness, or the teacher cannot do it) before starting it again.
  • DDLS: 8 of 17 runs accepted on 2026-10-05; the rejections were budget (60 calls, now 100) and empty responses (32k output limit). If empty responses stay high, lower the output limit for CDS runs or retry once more.

5. Settings that stay

  • Trajectory workers 2 (a harness event sets 1); stop rules: acceptance under 50 % over the last 30 runs, the same harness error three times, the budget guard. Generation deadline 2026-10-10 18:00 has passed: set --gen-deadline and --traj-deadline on the controller command (for example --gen-deadline 2026-10-20T18:00 --traj-deadline 2026-10-21T23:30).
  • Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
  • Summary every 50 accepted trajectories in train/STATE.md and docs/yol-haritasi.md, with a commit.

6. After the restart

Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the reserve), the EPOD requests in docs/epod-lock-leak.md and docs/epod-syntax-hint.md.