7.2 KiB
Handover notes (2026-10-03, 22:40; baseline shortened, see train/STATE.md)
Read this file, CLAUDE.md, train/STATE.md and train/README.md first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file.
0. Rule for the next session (Kral, 2026-10-03)
Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue
reruns for after the baseline in runs/stage1/rerun_queue.txt (one task id per line, with the reason).
The night chain still starts W7B and then the baseline as planned.
1. Running jobs
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
| Job | Command | PID | Log | Expected end |
|---|---|---|---|---|
| MLX server (base model, no adapter) | train/serve.sh |
59352 | runs/stage1/server.log |
runs until stopped (kill 59352) |
| T01 test run (budget 40, harness check only) | python3 train/baseline.py --label baseline --only T01 |
59412 (sh), 59414 (python) | runs/stage1/baseline_t01.log |
about 19:00 |
| DeepSeek filter W6A (last task G0125) | xargs python3 -m harness.empirical ... --run-base 15000 |
57246 / 57253 | runs/emp_deepseek6a.log |
about 18:45 |
| DeepSeek filter W7A (11 tasks, budget-60 reruns) | xargs ... --run-base 17000 (list runs/stage1/w7a.txt) |
59503 / 59510 | runs/emp_deepseek7a.log |
about 19:45 |
| Night chain | train/night_chain.sh |
60017 | runs/stage1/night_chain.log |
see below |
| Server restart with prompt cache limit (after the T01 test; new server PID in the log) | train/restart_server.sh |
60447 | runs/stage1/restart_server.log |
after the T01 test |
Night chain (train/night_chain.sh):
- When W6A ends: starts W7B (13 DeepSeek reruns, list
runs/stage1/w7b.txt, logruns/emp_deepseek7b.log, run base 18000). Expected end about 20:00. - When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks
T01 in
runs/stage1/baseline.json. If there is no harness error, it moves the T01 entry tot01_test_budget40and starts the baseline on all 25 subset tasks (budget 60):python3 train/baseline.py --label baseline --run-base 20100, logruns/stage1/baseline.log. If T01 failed: notification "T01 test failed: baseline not started", and it stops. - Notification "Baseline (11 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October morning to noon.
After the jobs:
- Read only the last lines of
runs/stage1/night_chain.logandruns/stage1/baseline.log. - Then stage 1 prompt B2 (
docs/stage1-step-prompts.md): write the summary fromruns/stage1/baseline.json. - Stop the MLX server before training (step 3); A4H must be stopped by Kral.
2. Open decisions
| Item | Status |
|---|---|
| G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. |
| Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. |
EPOD server bug: a second write of a PROG/FUNC is not activated (outcome: notExecuted), also not by sap_activate. |
Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. |
| Separate SAP user for the harness (dumps under user KESELI) | Decided 2026-10-03 (Kral): not needed. |
| G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). |
| A4H move to the MacBook (Podman) | Decided 2026-10-03 (Kral): not needed (2-3 h, risk). Instead set a Docker memory limit for A4H after the baseline (container restart; ask Kral first). |
3. Pending follow-ups
Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in
docs/faz1-tasarim.md 11h and docs/eval-inceleme.md.
Done:
- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2. G0107 (old G2 failure) is in W7B.
- Item 2: K specs checked, done.
- Item 5, review part: the 10 unreviewed tasks are reviewed and
accept. Low scores: G0108 model error, G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop. - Item 6: 10 random high-score tasks (seed 20261003) reviewed, all
accept. - Item 7:
tasks_gen/eval/easy_candidates.json(39 tasks; DeepSeek >= 95 and <= 20 tool calls). Candidates only; confirm with the Qwen baseline. - Item 8: written to
docs/faz1-tasarim.md11h anddocs/eval-inceleme.md.
Waiting for jobs (do these in the next session):
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
- After "Baseline (11 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the Kral spot-check list (flagged tasks + one per category, ~10).
- Rerun queue
runs/stage1/rerun_queue.txt(G0119, G0162): run after the baseline. Cost: DeepSeek, about 0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting. - A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (11 tasks) ended"; no
second question needed): check
docker statsand the HANA allocation limit first; choose the limit with the data (about 32 GB is the SAP minimum for the trial image; too low gives an OOM kill); clean stop ofa4h(give HANA time), apply the limit, start, testlocalhost:50000and one MCP call. The Mac mini has 64 GB; swap was 9 of 10 GB during the Qwen run. - Proxy finding:
sap_inactive_objectsshows objects of other runs (no prefix filter).Z0FFK001_IBAN_VALIDATORwas the object of the running T01 test (run prefix Z0FFK001), not a leftover: teardown removes it. Fix the proxy after the baseline; no A4H cleanup needed (checkcleanup-listafter the T01 test).
4. Stage 1 status
Done (step 0): train/.venv (Python 3.11.17 with uv), mlx-lm 0.32.0, weights
mlx-community/Qwen3.8-27B-4bit in ~/models/Qwen3.8-27B-4bit, train/serve.sh with the Ollama settings,
tool-call test passed, fixed subset train/subset.json, runner train/baseline.py, train/README.md.
Running: baseline (night chain). Not done: step 1 (data preparation, train/prepare.py); the valid loss of the
base model (step 2) needs step 1. Details and next steps: train/STATE.md.
5. Budget
- Kral pays 20 USD per month for 60 USD of DeepSeek usage (real money = usage / 3).
- Real usage (Ollama page, Kral, 2026-10-03 20:15): 24.63 of 60 USD. Ledger (list prices): 66.97 USD. New ratio: real ≈ ledger / 2.7 (the old ratio 1.47 was wrong). About 35 USD real left until 12 October.
- Guard:
BUDGET_LIMIT_USDraised from 75 to 135 (ledger) on 2026-10-03 (Kral): real cap about 50 of 60 USD. - Planned cloud work is small: reruns (G0119, G0162, ≈ 0.15 USD real) and judge calls for H tasks in the baseline. The baseline itself is local and free.
- The agent limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of the run cost.