7.1 KiB
Stage 1 state
Task: docs/stage1-training-task.md. Settings and weights: train/README.md. Updated 2026-10-03 22:40.
Done
-
Step 0 (setup):
train/.venv(Python 3.11.17, uv),mlx-lm0.32.0 /mlx0.32.3. Modelmlx-community/Qwen3.8-27B-4bitat~/models/Qwen3.8-27B-4bit(affine 4 bit, group 64, 16.1 GB; baseQwen/Qwen3.8-27B). The Ollama NVFP4 weights do not load in mlx_lm (global scale). -
Server:
train/serve.sh(mlx_lm.server, port 8080, thinking on withreasoning_effortmedium, temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed. -
Eval subset (Kral decision):
train/subset.json(copyruns/stage1/subset.json), 25 tasks. The task document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after training. -
Runner:
train/baseline.py --label <label>(one task at a time, resultsruns/stage1/<label>.json, run directoriesruns/stage1/<label>/). -
Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule):
train/prepare.py, reporttrain/data/report.md. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens. Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed. Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82 pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834; no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %. Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003). 748 iterations for 2 epochs (the count is the same as with the first dedup rule by coincidence).test.jsonl= copy of valid.
Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; runs/archive/devstral/). Official baseline: runs/stage1/baseline.json (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in
pgrep -f mlx_lm), logruns/stage1/server.log. - Baseline on the 11-task subset (
train/subset.json), thinking off, loop guard 3, max_tokens 16384, started 2026-10-04 07:59 bytrain/baseline_chain.sh(run base 20500; the guard also ends read loops), logsruns/stage1/baseline_chain.logandruns/stage1/baseline.log. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended"). - Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (
_aborted_*inruns/stage1/baseline/, A4H objects deleted).
Next
-
After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
-
B2: read the last lines of
runs/stage1/baseline.log; summary fromruns/stage1/baseline.json(t01_test_budget40holds the T01 test result). Commit. -
Step 2, second part: valid loss of the base model (
mlx_lm.lora --testwithout adapter; check the options with--helpfirst). Add it toruns/stage1/baseline.json. -
Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first.
Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit) is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(
max_tokensonly for cloud models; the MLX server limit is--max-tokens 32768). - Session 2026-10-03 (evening): follow-ups of
docs/devir-notlari.mdsection 3 done from stored results (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait inruns/stage1/rerun_queue.txt(G0119, G0162) until the baseline ends.
Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No mlx_lm training.
- Base valid loss (Mac, MLX 4-bit, no adapter): 0.849, ppl 2.337 (
runs/stage1/baseline.json, keyvalid_loss). - Done: private dataset
erhankeseli/abap-stage1-data; private model repoerhankeseli/abap-stage1-adapter-test;train/hf_train.py(Unsloth job),train/peft_to_mlx.py(converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters. - Key names: PEFT
base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight(A: r x in) -> mlxlanguage_model.model.layers.N.<mod>.lora_a/b(transposed). mlx scale = alpha / r. - Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx
scale: 20would be alpha 320. - Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5): 304.5 s train = 30.5 s/step, job wall time 8 min 7 s = about 0.34 USD, peak GPU 45.4 GB. Adapter pushed to the private repo. Unsloth valid loss (bnb 4-bit): 0.920.
- Conversion test:
peft_to_mlx.pyconverted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors. Valid loss on the Mac with the adapter: 0.848 (base 0.849,runs/stage1/adapter_test_valid_loss.log). Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same. Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded). Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps). - Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length). Not started. Needs Kral's go and the alpha decision.
Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)
- Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18), alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
- Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
- Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30,
timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line
EVAL), checkpoints and valid loss every 200 steps inerhankeseli/abap-stage1-adapter-full/step<N>/(witheval.json). Expected about 6.3 h. - Rule: at step 200 convert the checkpoint (
train/peft_to_mlx.py), measure the valid loss on the Mac (runs/stage1/adapter_test_valid_loss.logcommand, about 20 min), compare the relative drop with the GPU value. Stop the job (hf jobs cancel) if they differ.