Files
abap-llm/train/STATE.md

3.4 KiB

Stage 1 state

Task: docs/stage1-training-task.md. Settings and weights: train/README.md. Updated 2026-10-03 22:30.

Done

  • Step 0 (setup): train/.venv (Python 3.11.17, uv), mlx-lm 0.32.0 / mlx 0.32.3. Model mlx-community/Qwen3.8-27B-4bit at ~/models/Qwen3.8-27B-4bit (affine 4 bit, group 64, 16.1 GB; base Qwen/Qwen3.8-27B). The Ollama NVFP4 weights do not load in mlx_lm (global scale).

  • Server: train/serve.sh (mlx_lm.server, port 8080, thinking on with reasoning_effort medium, temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.

  • Eval subset (Kral decision): train/subset.json (copy runs/stage1/subset.json), 25 tasks. The task document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after training.

  • Runner: train/baseline.py --label <label> (one task at a time, results runs/stage1/<label>.json, run directories runs/stage1/<label>/).

  • Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): train/prepare.py, report train/data/report.md. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens. Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept; tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces (classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %. Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003). 618 -> 748 iterations for 2 epochs. test.jsonl = copy of valid.

Running (detached)

  • MLX server PID 59352, log runs/stage1/server.log.
  • T01 test (budget 40, harness check), PID 59414, log runs/stage1/baseline_t01.log.
  • Night chain train/night_chain.sh PID 60017, log runs/stage1/night_chain.log: after the DeepSeek reruns and the T01 test, it starts the baseline on the 25 tasks (budget 60), log runs/stage1/baseline.log, macOS notification at the end. Expected end: 4 October, morning to noon.

Next

  1. B2: read the last lines of runs/stage1/baseline.log; summary from runs/stage1/baseline.json (t01_test_budget40 holds the T01 test result). Commit.
  2. Step 2, second part: valid loss of the base model (mlx_lm.lora --test, no adapter) after the baseline.
  3. Step 2, second part: valid loss of the base model (mlx_lm.lora --test without adapter; check the options with --help first). Add it to runs/stage1/baseline.json.
  4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first.

Notes

  • The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit) is the new reference.
  • Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent (max_tokens only for cloud models; the MLX server limit is --max-tokens 32768).
  • Session 2026-10-03 (evening): follow-ups of docs/devir-notlari.md section 3 done from stored results (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in runs/stage1/rerun_queue.txt (G0119, G0162) until the baseline ends.