Files
abap-llm/train/STATE.md

75 lines
5.1 KiB
Markdown

# Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:40.
## Done
- Step 0 (setup): `train/.venv` (Python 3.11.17, uv), `mlx-lm` 0.32.0 / `mlx` 0.32.3.
Model `mlx-community/Qwen3.8-27B-4bit` at `~/models/Qwen3.8-27B-4bit` (affine 4 bit, group 64, 16.1 GB;
base `Qwen/Qwen3.8-27B`). The Ollama NVFP4 weights do not load in mlx_lm (global scale).
- Server: `train/serve.sh` (`mlx_lm.server`, port 8080, thinking on with `reasoning_effort` medium,
temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.
- Eval subset (Kral decision): `train/subset.json` (copy `runs/stage1/subset.json`), 25 tasks. The task
document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after
training.
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule): `train/prepare.py`, report
`train/data/report.md`. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with
the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed.
Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82
pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834;
no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %.
Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
`test.jsonl` = copy of valid.
## Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
## Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
2026-10-04 07:59 by `train/baseline_chain.sh` (run base 20500; the guard also ends read loops), logs `runs/stage1/baseline_chain.log` and
`runs/stage1/baseline.log`. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended").
- Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (`_aborted_*` in `runs/stage1/baseline/`, A4H objects deleted).
## Next
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
iterations first.
## Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit)
is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started.
- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost.