Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
43
train/STATE.md
Normal file
43
train/STATE.md
Normal file
@@ -0,0 +1,43 @@
|
||||
# Stage 1 state
|
||||
|
||||
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 18:35.
|
||||
|
||||
## Done
|
||||
|
||||
- Step 0 (setup): `train/.venv` (Python 3.11.17, uv), `mlx-lm` 0.32.0 / `mlx` 0.32.3.
|
||||
Model `mlx-community/Qwen3.8-27B-4bit` at `~/models/Qwen3.8-27B-4bit` (affine 4 bit, group 64, 16.1 GB;
|
||||
base `Qwen/Qwen3.8-27B`). The Ollama NVFP4 weights do not load in mlx_lm (global scale).
|
||||
- Server: `train/serve.sh` (`mlx_lm.server`, port 8080, thinking on with `reasoning_effort` medium,
|
||||
temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.
|
||||
- Eval subset (Kral decision): `train/subset.json` (copy `runs/stage1/subset.json`), 25 tasks. The task
|
||||
document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after
|
||||
training.
|
||||
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
|
||||
run directories `runs/stage1/<label>/`).
|
||||
|
||||
## Running (detached)
|
||||
|
||||
- MLX server PID 59352, log `runs/stage1/server.log`.
|
||||
- T01 test (budget 40, harness check), PID 59414, log `runs/stage1/baseline_t01.log`.
|
||||
- Night chain `train/night_chain.sh` PID 60017, log `runs/stage1/night_chain.log`: after the DeepSeek reruns
|
||||
and the T01 test, it starts the baseline on the 25 tasks (budget 60), log `runs/stage1/baseline.log`,
|
||||
macOS notification at the end. Expected end: 4 October, morning to noon.
|
||||
|
||||
## Next
|
||||
|
||||
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
||||
(`t01_test_budget40` holds the T01 test result). Commit.
|
||||
2. Step 1: `train/prepare.py` (real token counts with the base model tokenizer, length filter 16384,
|
||||
95/5 split by document, `train/data/train.jsonl` and `valid.jsonl`, report). A4H is not needed.
|
||||
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
||||
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
||||
iterations first.
|
||||
|
||||
## Notes
|
||||
|
||||
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit)
|
||||
is the new reference.
|
||||
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
|
||||
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
|
||||
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
|
||||
34
train/night_chain.sh
Executable file
34
train/night_chain.sh
Executable file
@@ -0,0 +1,34 @@
|
||||
#!/bin/sh
|
||||
# Night chain (2026-10-03), started detached so it survives the Claude Code session.
|
||||
# 1. When filter worker W6A (PID 57253) ends: start W7B (DeepSeek reruns, list in runs/stage1/w7b.txt).
|
||||
# 2. When the T01 test (PID 59414) and all DeepSeek workers end: check T01 in runs/stage1/baseline.json.
|
||||
# No harness error -> remove the T01 entry (it ran with the old budget 40) and run all 25 subset tasks.
|
||||
# 3. macOS notification at each end. Log: runs/stage1/night_chain.log
|
||||
cd "$(dirname "$0")/.."
|
||||
log() { echo "$(date '+%F %T') $*"; }
|
||||
while kill -0 57253 2>/dev/null; do sleep 30; done
|
||||
log "W6A ended; start W7B"
|
||||
nohup xargs python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 18000 \
|
||||
< runs/stage1/w7b.txt > runs/emp_deepseek7b.log 2>&1 &
|
||||
W7B=$!
|
||||
log "W7B xargs PID $W7B"
|
||||
while kill -0 59414 2>/dev/null || kill -0 59510 2>/dev/null || kill -0 $W7B 2>/dev/null; do sleep 60; done
|
||||
log "T01 test and DeepSeek workers ended"
|
||||
osascript -e 'display notification "DeepSeek reruns ended" with title "stage1"'
|
||||
OK=$(python3 -c "
|
||||
import json
|
||||
t=json.load(open('runs/stage1/baseline.json'))['tasks'].get('T01',{})
|
||||
print('yes' if t and 'error' not in t and t.get('score') is not None else 'no')")
|
||||
log "T01 test result usable: $OK"
|
||||
if [ "$OK" != "yes" ]; then
|
||||
osascript -e 'display notification "T01 test failed: baseline not started" with title "stage1"'
|
||||
log "baseline NOT started"; exit 1
|
||||
fi
|
||||
python3 - <<'PY'
|
||||
import json
|
||||
p='runs/stage1/baseline.json'; r=json.load(open(p)); r.setdefault('t01_test_budget40', r['tasks'].pop('T01')); json.dump(r,open(p,'w'),indent=1)
|
||||
PY
|
||||
log "start baseline, 25 tasks"
|
||||
python3 train/baseline.py --label baseline --run-base 20100 > runs/stage1/baseline.log 2>&1
|
||||
log "baseline ended: $(python3 -c "import json;print(len(json.load(open('runs/stage1/baseline.json'))['tasks']))") tasks"
|
||||
osascript -e 'display notification "Baseline (25 tasks) ended" with title "stage1"'
|
||||
Reference in New Issue
Block a user