Files
abap-llm/docs/restart-12-oktober.md

7.4 KiB

Restart plan for 12 October 2026 (cloud work)

Decision (Kral + Opus, 2026-10-05): when the budget guard stops, no cloud work (generation, trajectories) before the Ollama reset on 12 October. Until then only no-cloud work. This is the plan for the restart.

1. Before the start (5 minutes, no cloud call)

  1. Panel value after the reset (usage USD). Then: python3 -m harness.restart_plan --panel <value> It prints the .env values and the state below with live numbers. Put them in .env: BUDGET_CYCLE_START=2026-10-12, BUDGET_LIMIT_USD=<printed>, BUDGET_RESERVE_USD=8 (the 3 usage reserve for the second teacher test stays). Ratio for the guard: 1.2 ledger per usage (trajectory runs; generation is about 2, so the guard is on the safe side). Give the panel value again after about 2 hours and recompute (python3 -m harness.dashboard panel <v>).
  2. A4H up (docker ps), MCP answers, no stale lock: python3 scripts_probe/lockprobe.py must print 0 entries.
  3. One controller only: python3 -m harness.pipeline refuses to start a second one (runs/pipeline/controller.lock). Remove runs/pipeline/STOP and STOPPED.txt if they exist.
  4. Dashboard: python3 -m harness.dashboard (5-minute page, runs/dashboard/index.html).

1b. Step 0: eval tasks for the new kinds (before any training task)

python3 -m harness.evalset plan-new shows 32 slots (INTF 10, TABL 11, STRU 3, MSAG 4, exception 4; ids G0200 to G0231), then python3 -m harness.evalset run-new (about 5 ledger, about 2 usage), then python3 -m harness.evalset run-new-k (5 K variants). Reason: the eval set (130 accepted candidates) has no INTF, TABL, STRU, MSAG or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: docs/eval-spotcheck-new-kinds.md (Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.

1c. Step 0b: rerun of four empirical filter runs (Kral + Opus 2026-10-06)

Four eval candidates saw or read leftover objects of other runs in the first empirical filter (docs/foreign-objects-report.md): G0183, G0180, G0143, G0185 (DeepSeek, scores 80, 85, 85, 85). After step 0 (eval generation for the new kinds), with the fixed proxy:

python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 450000 --rerun G0183 G0180 G0143 G0185

This writes only to runs/emp_rerun/ (4 runs, about 1 ledger) and prints old and new filter decision per task. empirical.json and the eval set are not touched. Decision rules (step F): easy candidate = score 95 or more and at most 20 tool calls; flag = score under 50 (a strong model fails although the reference passes); else normal. If a decision changes for one of the four, report it to Kral and Opus first; the eval set is changed only after their answer. If nothing changes, say so in docs/eval-inceleme.md and leave the old results.

2. Order of work (what the controller does by itself)

Only kinds below their target share are generated and run (harness/mix.py, below_target); the kind with the biggest deficit goes first. Target (percent of accepted tasks and of accepted trajectories): CLAS 28, INTF 7, DDLS 25, FUNC 15, PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5.

  1. Trajectories for the new-type tasks that wait (INTF, TABL, STRU, MSAG, exception, PROG), then DDLS.
  2. Second attempts: a task whose first attempt failed, or was accepted without a repair, gets one more attempt (at most two attempts, at most two accepted trajectories per task). Failed DDLS and failed new-type tasks come first because they are below target.
  3. Generation (3 workers): the kind with the biggest deficit; a kind with more than 8 waiting tasks is not generated; a kind with 6 or more tries and under 20 % accepted is skipped. 20 % error-targeted slots, 30 % hard slots.
  4. CLAS and FUNC generation and trajectories wait until they are at or below their target share (CLAS is far above).
  5. K variants (free text) resume at 10 % of the other accepted tasks when their kinds are below target.

3. How much is needed (numbers of 2026-10-05 20:00, live: restart_plan)

CLAS has 49 accepted trajectories. At a 28 % share that is a total of 175 accepted trajectories, so about 105 more are needed, all on other kinds: INTF 12, DDLS about 35, FUNC about 15, PROG about 17, TABL 14, STRU 3, MSAG 4, exception 4. Tasks needed (accepted, with about 1.3 trajectory runs per accepted trajectory): TABL +9, INTF +10, STRU +3, MSAG +4, DDLS +20, PROG +5. About 130 trajectory runs at 0.25 ledger = 33 ledger = 27 usage, plus the generation (about 6 usage). Second attempts are part of the 130 runs. If the new reset gives 60 usage, this fits with room for the second teacher test.

4. Checks on the first runs of each new type (do not skip)

  • The first 5 runs of INTF, TABL, STRU, MSAG, exception: records complete (the proxy must pass sap_push_message), Qwen conversion works (train/to_qwen.py on accepted.jsonl: tool call round trip), no harness event.
  • Acceptance per kind in the summary; a kind under 20 % after 6 tries is skipped automatically: read why (prompt, harness, or the teacher cannot do it) before starting it again.
  • DDLS: 8 of 17 runs accepted on 2026-10-05; the rejections were budget (60 calls, now 100) and empty responses (32k output limit). If empty responses stay high, lower the output limit for CDS runs or retry once more.

5. Settings that stay

  • Trajectory workers 2 (a harness event sets 1); stop rules: acceptance under 50 % over the last 30 runs, the same harness error three times, the budget guard. Generation deadline 2026-10-10 18:00 has passed: set --gen-deadline and --traj-deadline on the controller command (for example --gen-deadline 2026-10-20T18:00 --traj-deadline 2026-10-21T23:30).
  • Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
  • Summary every 50 accepted trajectories in train/STATE.md and docs/yol-haritasi.md, with a commit.

5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks)

  1. git status clean and pushed; python3 -m harness.restart_plan --panel <value> prints the numbers (panel value from Kral, expected 0 after the reset).
  2. Services: A4H up, MCP answers, python3 scripts_probe/lockprobe.py shows no stale lock (only ATC runtime and debugger listener entries), python3 -m harness.sweep shows no leftover.
  3. .env: BUDGET_CYCLE_START=2026-10-12, BUDGET_LIMIT_USD, BUDGET_RESERVE_USD=8; no STOP flag, no STOPPED.txt; no running controller (runs/pipeline/controller.lock).
  4. Order of the first hour: step 0 eval slots for the new kinds (evalset plan-new, run-new, run-new-k), then the pending tasks of the kinds below target. The 9 trajectories that were moved back (runs/traj/summary_excluded.jsonl: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order.
  5. Settings to confirm: workers 2, STREAM_GUARD unset for the first runs and then 9000 for 5 DDLS/PROG runs (docs/empty-response.md), tool budget 100 for CDS.
  6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host.

6. After the restart

Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the reserve), the EPOD requests in docs/epod-lock-leak.md and docs/epod-syntax-hint.md.