Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2.5 KiB
2.5 KiB
Stage 1 state
Task: docs/stage1-training-task.md. Settings and weights: train/README.md. Updated 2026-10-03 18:35.
Done
- Step 0 (setup):
train/.venv(Python 3.11.17, uv),mlx-lm0.32.0 /mlx0.32.3. Modelmlx-community/Qwen3.8-27B-4bitat~/models/Qwen3.8-27B-4bit(affine 4 bit, group 64, 16.1 GB; baseQwen/Qwen3.8-27B). The Ollama NVFP4 weights do not load in mlx_lm (global scale). - Server:
train/serve.sh(mlx_lm.server, port 8080, thinking on withreasoning_effortmedium, temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed. - Eval subset (Kral decision):
train/subset.json(copyruns/stage1/subset.json), 25 tasks. The task document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after training. - Runner:
train/baseline.py --label <label>(one task at a time, resultsruns/stage1/<label>.json, run directoriesruns/stage1/<label>/).
Running (detached)
- MLX server PID 59352, log
runs/stage1/server.log. - T01 test (budget 40, harness check), PID 59414, log
runs/stage1/baseline_t01.log. - Night chain
train/night_chain.shPID 60017, logruns/stage1/night_chain.log: after the DeepSeek reruns and the T01 test, it starts the baseline on the 25 tasks (budget 60), logruns/stage1/baseline.log, macOS notification at the end. Expected end: 4 October, morning to noon.
Next
- B2: read the last lines of
runs/stage1/baseline.log; summary fromruns/stage1/baseline.json(t01_test_budget40holds the T01 test result). Commit. - Step 1:
train/prepare.py(real token counts with the base model tokenizer, length filter 16384, 95/5 split by document,train/data/train.jsonlandvalid.jsonl, report). A4H is not needed. - Step 2, second part: valid loss of the base model (
mlx_lm.lora --testwithout adapter; check the options with--helpfirst). Add it toruns/stage1/baseline.json. - Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first.
Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit) is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(
max_tokensonly for cloud models; the MLX server limit is--max-tokens 32768).