# Handover notes (2026-10-03, 19:10) Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into CLAUDE.md and the docs when they are done, then empty this file. ## 0. Rule for the next session (Kral, 2026-10-03) Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue reruns for after the baseline in `runs/stage1/rerun_queue.txt` (one task id per line, with the reason). The night chain still starts W7B and then the baseline as planned. ## 1. Running jobs All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends. | Job | Command | PID | Log | Expected end | |---|---|---|---|---| | MLX server (base model, no adapter) | `train/serve.sh` | 59352 | `runs/stage1/server.log` | runs until stopped (`kill 59352`) | | T01 test run (budget 40, harness check only) | `python3 train/baseline.py --label baseline --only T01` | 59412 (sh), 59414 (python) | `runs/stage1/baseline_t01.log` | about 19:00 | | DeepSeek filter W6A (last task G0125) | `xargs python3 -m harness.empirical ... --run-base 15000` | 57246 / 57253 | `runs/emp_deepseek6a.log` | about 18:45 | | DeepSeek filter W7A (11 tasks, budget-60 reruns) | `xargs ... --run-base 17000 (list runs/stage1/w7a.txt)` | 59503 / 59510 | `runs/emp_deepseek7a.log` | about 19:45 | | Night chain | `train/night_chain.sh` | 60017 | `runs/stage1/night_chain.log` | see below | | Server restart with prompt cache limit (after the T01 test; new server PID in the log) | `train/restart_server.sh` | 60447 | `runs/stage1/restart_server.log` | after the T01 test | Night chain (`train/night_chain.sh`): 1. When W6A ends: starts W7B (13 DeepSeek reruns, list `runs/stage1/w7b.txt`, log `runs/emp_deepseek7b.log`, run base 18000). Expected end about 20:00. 2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks T01 in `runs/stage1/baseline.json`. If there is no harness error, it moves the T01 entry to `t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60): `python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`. If T01 failed: notification "T01 test failed: baseline not started", and it stops. 3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October morning to noon. After the jobs: - Read only the last lines of `runs/stage1/night_chain.log` and `runs/stage1/baseline.log`. - Then stage 1 prompt B2 (`docs/stage1-step-prompts.md`): write the summary from `runs/stage1/baseline.json`. - Stop the MLX server before training (step 3); A4H must be stopped by Kral. ## 2. Open decisions | Item | Status | |---|---| | G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. | | Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. | | EPOD server bug: a second write of a PROG/FUNC is not activated (`outcome: notExecuted`), also not by `sap_activate`. | Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. | | Separate SAP user for the harness (dumps under user KESELI) | Decided 2026-10-03 (Kral): not needed. | | G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). | | A4H move to the MacBook (Podman) | Decided 2026-10-03 (Kral): not needed (2-3 h, risk). Instead set a Docker memory limit for A4H after the baseline (container restart; ask Kral first). | ## 3. Pending follow-ups Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`. Done: - Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2. G0107 (old G2 failure) is in W7B. - Item 2: K specs checked, done. - Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error, G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop. - Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`. - Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls). Candidates only; confirm with the Qwen baseline. - Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`. Waiting for jobs (do these in the next session): - After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?). - After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the Kral spot-check list (flagged tasks + one per category, ~10). - Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about 0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting. - A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (25 tasks) ended"; no second question needed): check `docker stats` and the HANA allocation limit first; choose the limit with the data (about 32 GB is the SAP minimum for the trial image; too low gives an OOM kill); clean stop of `a4h` (give HANA time), apply the limit, start, test `localhost:50000` and one MCP call. The Mac mini has 64 GB; swap was 9 of 10 GB during the Qwen run. - Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter). `Z0FFK001_IBAN_VALIDATOR` was the object of the running T01 test (run prefix Z0FFK001), not a leftover: teardown removes it. Fix the proxy after the baseline; no A4H cleanup needed (check `cleanup-list` after the T01 test). ## 4. Stage 1 status Done (step 0): `train/.venv` (Python 3.11.17 with uv), `mlx-lm` 0.32.0, weights `mlx-community/Qwen3.8-27B-4bit` in `~/models/Qwen3.8-27B-4bit`, `train/serve.sh` with the Ollama settings, tool-call test passed, fixed subset `train/subset.json`, runner `train/baseline.py`, `train/README.md`. Running: baseline (night chain). Not done: step 1 (data preparation, `train/prepare.py`); the valid loss of the base model (step 2) needs step 1. Details and next steps: `train/STATE.md`. ## 5. Budget - Real usage (Ollama page, Kral, 2026-10-03 20:15): 24.63 of 60 USD. Ledger (list prices): 66.97 USD. New ratio: real ≈ ledger / 2.7 (the old ratio 1.47 was wrong). About 35 USD real left until 12 October. - Guard: `BUDGET_LIMIT_USD=75` (ledger) is now about 28 USD real: too strict. Proposal: raise it so that the real cap is about 50 of 60 USD: ledger ≈ 135. Not changed yet; Kral decides. - Planned cloud work is small: reruns (G0119, G0162, ≈ 0.15 USD real) and judge calls for H tasks in the baseline. The baseline itself is local and free. - The agent limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of the run cost.