94 lines
6.2 KiB
Markdown
94 lines
6.2 KiB
Markdown
# Handover notes (2026-10-03, 19:10)
|
||
|
||
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
|
||
CLAUDE.md and the docs when they are done, then empty this file.
|
||
|
||
## 0. Rule for the next session (Kral, 2026-10-03)
|
||
|
||
Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue
|
||
reruns for after the baseline in `runs/stage1/rerun_queue.txt` (one task id per line, with the reason).
|
||
The night chain still starts W7B and then the baseline as planned.
|
||
|
||
## 1. Running jobs
|
||
|
||
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
|
||
|
||
| Job | Command | PID | Log | Expected end |
|
||
|---|---|---|---|---|
|
||
| MLX server (base model, no adapter) | `train/serve.sh` | 59352 | `runs/stage1/server.log` | runs until stopped (`kill 59352`) |
|
||
| T01 test run (budget 40, harness check only) | `python3 train/baseline.py --label baseline --only T01` | 59412 (sh), 59414 (python) | `runs/stage1/baseline_t01.log` | about 19:00 |
|
||
| DeepSeek filter W6A (last task G0125) | `xargs python3 -m harness.empirical ... --run-base 15000` | 57246 / 57253 | `runs/emp_deepseek6a.log` | about 18:45 |
|
||
| DeepSeek filter W7A (11 tasks, budget-60 reruns) | `xargs ... --run-base 17000 (list runs/stage1/w7a.txt)` | 59503 / 59510 | `runs/emp_deepseek7a.log` | about 19:45 |
|
||
| Night chain | `train/night_chain.sh` | 60017 | `runs/stage1/night_chain.log` | see below |
|
||
| Server restart with prompt cache limit (after the T01 test; new server PID in the log) | `train/restart_server.sh` | 60447 | `runs/stage1/restart_server.log` | after the T01 test |
|
||
|
||
Night chain (`train/night_chain.sh`):
|
||
1. When W6A ends: starts W7B (13 DeepSeek reruns, list `runs/stage1/w7b.txt`, log `runs/emp_deepseek7b.log`,
|
||
run base 18000). Expected end about 20:00.
|
||
2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks
|
||
T01 in `runs/stage1/baseline.json`. If there is no harness error, it moves the T01 entry to
|
||
`t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60):
|
||
`python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`.
|
||
If T01 failed: notification "T01 test failed: baseline not started", and it stops.
|
||
3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October
|
||
morning to noon.
|
||
|
||
After the jobs:
|
||
- Read only the last lines of `runs/stage1/night_chain.log` and `runs/stage1/baseline.log`.
|
||
- Then stage 1 prompt B2 (`docs/stage1-step-prompts.md`): write the summary from `runs/stage1/baseline.json`.
|
||
- Stop the MLX server before training (step 3); A4H must be stopped by Kral.
|
||
|
||
## 2. Open decisions
|
||
|
||
| Item | Status |
|
||
|---|---|
|
||
| G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. |
|
||
| Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. |
|
||
| EPOD server bug: a second write of a PROG/FUNC is not activated (`outcome: notExecuted`), also not by `sap_activate`. | Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. |
|
||
| Separate SAP user for the harness (dumps under user KESELI) | Decided 2026-10-03 (Kral): not needed. |
|
||
| G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). |
|
||
|
||
## 3. Pending follow-ups
|
||
|
||
Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in
|
||
`docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
|
||
|
||
Done:
|
||
- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2.
|
||
G0107 (old G2 failure) is in W7B.
|
||
- Item 2: K specs checked, done.
|
||
- Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error,
|
||
G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop.
|
||
- Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`.
|
||
- Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls).
|
||
Candidates only; confirm with the Qwen baseline.
|
||
- Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
|
||
|
||
Waiting for jobs (do these in the next session):
|
||
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and
|
||
the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
|
||
- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
|
||
Kral spot-check list (flagged tasks + one per category, ~10).
|
||
- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about
|
||
0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
|
||
- Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter); an inactive object
|
||
`Z0FFK001_IBAN_VALIDATOR` stays on A4H. Fix the proxy and clean up after the baseline (A4H change: ask Kral).
|
||
|
||
## 4. Stage 1 status
|
||
|
||
Done (step 0): `train/.venv` (Python 3.11.17 with uv), `mlx-lm` 0.32.0, weights
|
||
`mlx-community/Qwen3.8-27B-4bit` in `~/models/Qwen3.8-27B-4bit`, `train/serve.sh` with the Ollama settings,
|
||
tool-call test passed, fixed subset `train/subset.json`, runner `train/baseline.py`, `train/README.md`.
|
||
Running: baseline (night chain). Not done: step 1 (data preparation, `train/prepare.py`); the valid loss of the
|
||
base model (step 2) needs step 1. Details and next steps: `train/STATE.md`.
|
||
|
||
## 5. Budget
|
||
|
||
- Cycle budget 60 USD until 12 October 2026. Ledger 63.3 USD (list prices) ≈ 43 USD real (ratio 1.47,
|
||
measured 2026-10-03). About 17 USD real left.
|
||
- Guard: `BUDGET_LIMIT_USD=75` (ledger) ≈ 51 USD real.
|
||
- Tonight: about 25 DeepSeek reruns (≈ 1.5 USD real) and judge calls for H tasks in the baseline (small).
|
||
The baseline itself is local and free.
|
||
- The agent now limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of
|
||
the run cost.
|