Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
55 lines
3.4 KiB
Markdown
55 lines
3.4 KiB
Markdown
# Stage 1 state
|
|
|
|
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:30.
|
|
|
|
## Done
|
|
|
|
- Step 0 (setup): `train/.venv` (Python 3.11.17, uv), `mlx-lm` 0.32.0 / `mlx` 0.32.3.
|
|
Model `mlx-community/Qwen3.8-27B-4bit` at `~/models/Qwen3.8-27B-4bit` (affine 4 bit, group 64, 16.1 GB;
|
|
base `Qwen/Qwen3.8-27B`). The Ollama NVFP4 weights do not load in mlx_lm (global scale).
|
|
- Server: `train/serve.sh` (`mlx_lm.server`, port 8080, thinking on with `reasoning_effort` medium,
|
|
temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.
|
|
- Eval subset (Kral decision): `train/subset.json` (copy `runs/stage1/subset.json`), 25 tasks. The task
|
|
document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after
|
|
training.
|
|
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
|
|
run directories `runs/stage1/<label>/`).
|
|
|
|
- Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): `train/prepare.py`, report `train/data/report.md`.
|
|
Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
|
|
Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept;
|
|
tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces
|
|
(classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece
|
|
over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %.
|
|
Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
|
|
**618 -> 748 iterations for 2 epochs.** `test.jsonl` = copy of valid.
|
|
|
|
## Running (detached)
|
|
|
|
- MLX server PID 59352, log `runs/stage1/server.log`.
|
|
- T01 test (budget 40, harness check), PID 59414, log `runs/stage1/baseline_t01.log`.
|
|
- Night chain `train/night_chain.sh` PID 60017, log `runs/stage1/night_chain.log`: after the DeepSeek reruns
|
|
and the T01 test, it starts the baseline on the 25 tasks (budget 60), log `runs/stage1/baseline.log`,
|
|
macOS notification at the end. Expected end: 4 October, morning to noon.
|
|
|
|
## Next
|
|
|
|
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
|
(`t01_test_budget40` holds the T01 test result). Commit.
|
|
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
|
|
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
|
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
|
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
|
iterations first.
|
|
|
|
## Notes
|
|
|
|
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit)
|
|
is the new reference.
|
|
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
|
|
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
|
|
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
|
|
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
|
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
|
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|