Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 18:34:07 +02:00
parent d4ed32cf38
commit 9366ee9868
7 changed files with 222 additions and 0 deletions

81
docs/devir-notlari.md Normal file
View File

@@ -0,0 +1,81 @@
# Handover notes (2026-10-03, 18:35)
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file.
## 1. Running jobs
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
| Job | Command | PID | Log | Expected end |
|---|---|---|---|---|
| MLX server (base model, no adapter) | `train/serve.sh` | 59352 | `runs/stage1/server.log` | runs until stopped (`kill 59352`) |
| T01 test run (budget 40, harness check only) | `python3 train/baseline.py --label baseline --only T01` | 59412 (sh), 59414 (python) | `runs/stage1/baseline_t01.log` | about 19:00 |
| DeepSeek filter W6A (last task G0125) | `xargs python3 -m harness.empirical ... --run-base 15000` | 57246 / 57253 | `runs/emp_deepseek6a.log` | about 18:45 |
| DeepSeek filter W7A (11 tasks, budget-60 reruns) | `xargs ... --run-base 17000 (list runs/stage1/w7a.txt)` | 59503 / 59510 | `runs/emp_deepseek7a.log` | about 19:45 |
| Night chain | `train/night_chain.sh` | 60017 | `runs/stage1/night_chain.log` | see below |
Night chain (`train/night_chain.sh`):
1. When W6A ends: starts W7B (13 DeepSeek reruns, list `runs/stage1/w7b.txt`, log `runs/emp_deepseek7b.log`,
run base 18000). Expected end about 20:00.
2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks
T01 in `runs/stage1/baseline.json`. If there is no harness error, it moves the T01 entry to
`t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60):
`python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`.
If T01 failed: notification "T01 test failed: baseline not started", and it stops.
3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October
morning to noon.
After the jobs:
- Read only the last lines of `runs/stage1/night_chain.log` and `runs/stage1/baseline.log`.
- Then stage 1 prompt B2 (`docs/stage1-step-prompts.md`): write the summary from `runs/stage1/baseline.json`.
- Stop the MLX server before training (step 3); A4H must be stopped by Kral.
## 2. Open decisions
| Item | Status |
|---|---|
| G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. |
| Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. |
| EPOD server bug: a second write of a PROG/FUNC is not activated (`outcome: notExecuted`), also not by `sap_activate`. | Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. |
| Separate SAP user for the harness (dumps under user KESELI) | Proposed, no decision. |
| G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). |
## 3. Pending follow-ups
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
G0107). Old DeepSeek runs of FUNC tasks where G2 failed: re-score or rerun. G0107 is queued in W7B.
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
response"). Not yet queued.
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
(G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189).
6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done.
7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few
tool calls; confirm with the Qwen baseline.
8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and
`docs/eval-inceleme.md`.
## 4. Stage 1 status
Done (step 0): `train/.venv` (Python 3.11.17 with uv), `mlx-lm` 0.32.0, weights
`mlx-community/Qwen3.8-27B-4bit` in `~/models/Qwen3.8-27B-4bit`, `train/serve.sh` with the Ollama settings,
tool-call test passed, fixed subset `train/subset.json`, runner `train/baseline.py`, `train/README.md`.
Running: baseline (night chain). Not done: step 1 (data preparation, `train/prepare.py`); the valid loss of the
base model (step 2) needs step 1. Details and next steps: `train/STATE.md`.
## 5. Budget
- Cycle budget 60 USD until 12 October 2026. Ledger 63.3 USD (list prices) ≈ 43 USD real (ratio 1.47,
measured 2026-10-03). About 17 USD real left.
- Guard: `BUDGET_LIMIT_USD=75` (ledger) ≈ 51 USD real.
- Tonight: about 25 DeepSeek reruns (≈ 1.5 USD real) and judge calls for H tasks in the baseline (small).
The baseline itself is local and free.
- The agent now limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of
the run cost.

View File

@@ -15,5 +15,21 @@
"seconds": 826.4,
"run_dir": "runs/emp/11008_G0123_llm_deepseek-v4.1-flash_cloud"
}
},
"deepseek-v4.1-flash:cloud": {
"score": 65.0,
"parts": {
"correctness": 40.0,
"craft": 5.0,
"clean": 15.0,
"own_tests": 0.0,
"discipline": 5.0,
"total": 65.0,
"note": "own_tests: faulty-reference part not implemented in skeleton"
},
"hidden": "12/12",
"tool_calls": 27,
"seconds": 565.3,
"run_dir": "runs/emp/15013_G0123_llm_deepseek-v4.1-flash_cloud"
}
}

View File

@@ -16,5 +16,21 @@
"seconds": 144.2,
"run_dir": "runs/emp/11025_G0155_llm_deepseek-v4.1-flash_cloud"
}
},
"deepseek-v4.1-flash:cloud": {
"score": 80.0,
"parts": {
"correctness": 40.0,
"craft": 20.0,
"clean": 15.0,
"own_tests": 0.0,
"discipline": 5.0,
"total": 80.0,
"note": "own_tests: faulty-reference part not implemented in skeleton"
},
"hidden": "10/10",
"tool_calls": 28,
"seconds": 380.6,
"run_dir": "runs/emp/17000_G0155_llm_deepseek-v4.1-flash_cloud"
}
}

View File

@@ -16,5 +16,21 @@
"seconds": 45.7,
"run_dir": "runs/emp/11027_G0157_llm_deepseek-v4.1-flash_cloud"
}
},
"deepseek-v4.1-flash:cloud": {
"score": 85.0,
"parts": {
"correctness": 40.0,
"craft": 20.0,
"clean": 15.0,
"own_tests": 0.0,
"discipline": 10.0,
"total": 85.0,
"note": "own_tests: faulty-reference part not implemented in skeleton"
},
"hidden": "10/10",
"tool_calls": 10,
"seconds": 64.6,
"run_dir": "runs/emp/17001_G0157_llm_deepseek-v4.1-flash_cloud"
}
}

View File

@@ -16,5 +16,21 @@
"seconds": 36.2,
"run_dir": "runs/emp/11029_G0159_llm_deepseek-v4.1-flash_cloud"
}
},
"deepseek-v4.1-flash:cloud": {
"score": 100.0,
"parts": {
"correctness": 40.0,
"craft": 20.0,
"clean": 15.0,
"own_tests": 15.0,
"discipline": 10.0,
"total": 100.0,
"note": "own_tests: faulty-reference part not implemented in skeleton"
},
"hidden": "10/10",
"tool_calls": 10,
"seconds": 62.0,
"run_dir": "runs/emp/17002_G0159_llm_deepseek-v4.1-flash_cloud"
}
}

43
train/STATE.md Normal file
View File

@@ -0,0 +1,43 @@
# Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 18:35.
## Done
- Step 0 (setup): `train/.venv` (Python 3.11.17, uv), `mlx-lm` 0.32.0 / `mlx` 0.32.3.
Model `mlx-community/Qwen3.8-27B-4bit` at `~/models/Qwen3.8-27B-4bit` (affine 4 bit, group 64, 16.1 GB;
base `Qwen/Qwen3.8-27B`). The Ollama NVFP4 weights do not load in mlx_lm (global scale).
- Server: `train/serve.sh` (`mlx_lm.server`, port 8080, thinking on with `reasoning_effort` medium,
temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.
- Eval subset (Kral decision): `train/subset.json` (copy `runs/stage1/subset.json`), 25 tasks. The task
document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after
training.
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`).
## Running (detached)
- MLX server PID 59352, log `runs/stage1/server.log`.
- T01 test (budget 40, harness check), PID 59414, log `runs/stage1/baseline_t01.log`.
- Night chain `train/night_chain.sh` PID 60017, log `runs/stage1/night_chain.log`: after the DeepSeek reruns
and the T01 test, it starts the baseline on the 25 tasks (budget 60), log `runs/stage1/baseline.log`,
macOS notification at the end. Expected end: 4 October, morning to noon.
## Next
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Step 1: `train/prepare.py` (real token counts with the base model tokenizer, length filter 16384,
95/5 split by document, `train/data/train.jsonl` and `valid.jsonl`, report). A4H is not needed.
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
iterations first.
## Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit)
is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).

34
train/night_chain.sh Executable file
View File

@@ -0,0 +1,34 @@
#!/bin/sh
# Night chain (2026-10-03), started detached so it survives the Claude Code session.
# 1. When filter worker W6A (PID 57253) ends: start W7B (DeepSeek reruns, list in runs/stage1/w7b.txt).
# 2. When the T01 test (PID 59414) and all DeepSeek workers end: check T01 in runs/stage1/baseline.json.
# No harness error -> remove the T01 entry (it ran with the old budget 40) and run all 25 subset tasks.
# 3. macOS notification at each end. Log: runs/stage1/night_chain.log
cd "$(dirname "$0")/.."
log() { echo "$(date '+%F %T') $*"; }
while kill -0 57253 2>/dev/null; do sleep 30; done
log "W6A ended; start W7B"
nohup xargs python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 18000 \
< runs/stage1/w7b.txt > runs/emp_deepseek7b.log 2>&1 &
W7B=$!
log "W7B xargs PID $W7B"
while kill -0 59414 2>/dev/null || kill -0 59510 2>/dev/null || kill -0 $W7B 2>/dev/null; do sleep 60; done
log "T01 test and DeepSeek workers ended"
osascript -e 'display notification "DeepSeek reruns ended" with title "stage1"'
OK=$(python3 -c "
import json
t=json.load(open('runs/stage1/baseline.json'))['tasks'].get('T01',{})
print('yes' if t and 'error' not in t and t.get('score') is not None else 'no')")
log "T01 test result usable: $OK"
if [ "$OK" != "yes" ]; then
osascript -e 'display notification "T01 test failed: baseline not started" with title "stage1"'
log "baseline NOT started"; exit 1
fi
python3 - <<'PY'
import json
p='runs/stage1/baseline.json'; r=json.load(open(p)); r.setdefault('t01_test_budget40', r['tasks'].pop('T01')); json.dump(r,open(p,'w'),indent=1)
PY
log "start baseline, 25 tasks"
python3 train/baseline.py --label baseline --run-base 20100 > runs/stage1/baseline.log 2>&1
log "baseline ended: $(python3 -c "import json;print(len(json.load(open('runs/stage1/baseline.json'))['tasks']))") tasks"
osascript -e 'display notification "Baseline (25 tasks) ended" with title "stage1"'