# Handover notes (2026-10-03, 18:35) Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into CLAUDE.md and the docs when they are done, then empty this file. ## 1. Running jobs All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends. | Job | Command | PID | Log | Expected end | |---|---|---|---|---| | MLX server (base model, no adapter) | `train/serve.sh` | 59352 | `runs/stage1/server.log` | runs until stopped (`kill 59352`) | | T01 test run (budget 40, harness check only) | `python3 train/baseline.py --label baseline --only T01` | 59412 (sh), 59414 (python) | `runs/stage1/baseline_t01.log` | about 19:00 | | DeepSeek filter W6A (last task G0125) | `xargs python3 -m harness.empirical ... --run-base 15000` | 57246 / 57253 | `runs/emp_deepseek6a.log` | about 18:45 | | DeepSeek filter W7A (11 tasks, budget-60 reruns) | `xargs ... --run-base 17000 (list runs/stage1/w7a.txt)` | 59503 / 59510 | `runs/emp_deepseek7a.log` | about 19:45 | | Night chain | `train/night_chain.sh` | 60017 | `runs/stage1/night_chain.log` | see below | Night chain (`train/night_chain.sh`): 1. When W6A ends: starts W7B (13 DeepSeek reruns, list `runs/stage1/w7b.txt`, log `runs/emp_deepseek7b.log`, run base 18000). Expected end about 20:00. 2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks T01 in `runs/stage1/baseline.json`. If there is no harness error, it moves the T01 entry to `t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60): `python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`. If T01 failed: notification "T01 test failed: baseline not started", and it stops. 3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October morning to noon. After the jobs: - Read only the last lines of `runs/stage1/night_chain.log` and `runs/stage1/baseline.log`. - Then stage 1 prompt B2 (`docs/stage1-step-prompts.md`): write the summary from `runs/stage1/baseline.json`. - Stop the MLX server before training (step 3); A4H must be stopped by Kral. ## 2. Open decisions | Item | Status | |---|---| | G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. | | Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. | | EPOD server bug: a second write of a PROG/FUNC is not activated (`outcome: notExecuted`), also not by `sap_activate`. | Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. | | Separate SAP user for the harness (dumps under user KESELI) | Proposed, no decision. | | G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). | ## 3. Pending follow-ups 1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes, G0107). Old DeepSeek runs of FUNC tasks where G2 failed: re-score or rerun. G0107 is queued in W7B. 2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form. 3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations (G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued (W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60. 4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model response"). Not yet queued. 5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open). To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks (G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189). 6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done. 7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few tool calls; confirm with the Qwen baseline. 8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and `docs/eval-inceleme.md`. ## 4. Stage 1 status Done (step 0): `train/.venv` (Python 3.11.17 with uv), `mlx-lm` 0.32.0, weights `mlx-community/Qwen3.8-27B-4bit` in `~/models/Qwen3.8-27B-4bit`, `train/serve.sh` with the Ollama settings, tool-call test passed, fixed subset `train/subset.json`, runner `train/baseline.py`, `train/README.md`. Running: baseline (night chain). Not done: step 1 (data preparation, `train/prepare.py`); the valid loss of the base model (step 2) needs step 1. Details and next steps: `train/STATE.md`. ## 5. Budget - Cycle budget 60 USD until 12 October 2026. Ledger 63.3 USD (list prices) ≈ 43 USD real (ratio 1.47, measured 2026-10-03). About 17 USD real left. - Guard: `BUDGET_LIMIT_USD=75` (ledger) ≈ 51 USD real. - Tonight: about 25 DeepSeek reruns (≈ 1.5 USD real) and judge calls for H tasks in the baseline (small). The baseline itself is local and free. - The agent now limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of the run cost.