Files
abap-llm/train/STATE.md
2026-10-04 20:48:49 +02:00

8.2 KiB

Stage 1 state

Task: docs/stage1-training-task.md. Settings and weights: train/README.md. Updated 2026-10-03 22:40.

Done

  • Step 0 (setup): train/.venv (Python 3.11.17, uv), mlx-lm 0.32.0 / mlx 0.32.3. Model mlx-community/Qwen3.8-27B-4bit at ~/models/Qwen3.8-27B-4bit (affine 4 bit, group 64, 16.1 GB; base Qwen/Qwen3.8-27B). The Ollama NVFP4 weights do not load in mlx_lm (global scale).

  • Server: train/serve.sh (mlx_lm.server, port 8080, thinking on with reasoning_effort medium, temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.

  • Eval subset (Kral decision): train/subset.json (copy runs/stage1/subset.json), 25 tasks. The task document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after training.

  • Runner: train/baseline.py --label <label> (one task at a time, results runs/stage1/<label>.json, run directories runs/stage1/<label>/).

  • Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule): train/prepare.py, report train/data/report.md. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens. Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed. Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82 pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834; no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %. Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003). 748 iterations for 2 epochs (the count is the same as with the first dedup rule by coincidence). test.jsonl = copy of valid.

Base model

Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; runs/archive/devstral/). Official baseline: runs/stage1/baseline.json (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.

Running (detached) — historical, all ended

  • MLX server (restarted 2026-10-04 06:38 after it had exited; PID in pgrep -f mlx_lm), log runs/stage1/server.log.
  • Baseline on the 11-task subset (train/subset.json), thinking off, loop guard 3, max_tokens 16384, started 2026-10-04 07:59 by train/baseline_chain.sh (run base 20500; the guard also ends read loops), logs runs/stage1/baseline_chain.log and runs/stage1/baseline.log. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended").
  • Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (_aborted_* in runs/stage1/baseline/, A4H objects deleted).

Next

  1. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.

  2. B2: read the last lines of runs/stage1/baseline.log; summary from runs/stage1/baseline.json (t01_test_budget40 holds the T01 test result). Commit.

  3. Step 2, second part: valid loss of the base model (mlx_lm.lora --test without adapter; check the options with --help first). Add it to runs/stage1/baseline.json.

  4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first.

Notes

  • The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit) is the new reference.
  • Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent (max_tokens only for cloud models; the MLX server limit is --max-tokens 32768).
  • Session 2026-10-03 (evening): follow-ups of docs/devir-notlari.md section 3 done from stored results (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in runs/stage1/rerun_queue.txt (G0119, G0162) until the baseline ends.

Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)

Training runs on HF Jobs with Unsloth, not on the Mac. No mlx_lm training.

  • Base valid loss (Mac, MLX 4-bit, no adapter): 0.849, ppl 2.337 (runs/stage1/baseline.json, key valid_loss).
  • Done: private dataset erhankeseli/abap-stage1-data; private model repo erhankeseli/abap-stage1-adapter-test; train/hf_train.py (Unsloth job), train/peft_to_mlx.py (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
  • Key names: PEFT base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight (A: r x in) -> mlx language_model.model.layers.N.<mod>.lora_a/b (transposed). mlx scale = alpha / r.
  • Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx scale: 20 would be alpha 320.
  • Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5): 304.5 s train = 30.5 s/step, job wall time 8 min 7 s = about 0.34 USD, peak GPU 45.4 GB. Adapter pushed to the private repo. Unsloth valid loss (bnb 4-bit): 0.920.
  • Conversion test: peft_to_mlx.py converted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors. Valid loss on the Mac with the adapter: 0.848 (base 0.849, runs/stage1/adapter_test_valid_loss.log). Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same. Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded). Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps).
  • Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length). Not started. Needs Kral's go and the alpha decision.

Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)

  • Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18), alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
  • Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
  • Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30, timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line EVAL), checkpoints and valid loss every 200 steps in erhankeseli/abap-stage1-adapter-full/step<N>/ (with eval.json). Expected about 6.3 h.
  • Rule: at step 200 convert the checkpoint (train/peft_to_mlx.py), measure the valid loss on the Mac (runs/stage1/adapter_test_valid_loss.log command, about 20 min), compare the relative drop with the GPU value. Stop the job (hf jobs cancel) if they differ.

Step 200 check failed, full run cancelled (2026-10-04)

  • GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
  • Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log runs/stage1/full_step200_valid_loss.log.
  • Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD). Checkpoints step200 and step400 stay in erhankeseli/abap-stage1-adapter-full.
  • Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4 error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091 on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
  • Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA, larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.