12 KiB
Stage 1 state
Task: docs/stage1-training-task.md. Settings and weights: train/README.md. Updated 2026-10-03 22:40.
Done
-
Step 0 (setup):
train/.venv(Python 3.11.17, uv),mlx-lm0.32.0 /mlx0.32.3. Modelmlx-community/Qwen3.8-27B-4bitat~/models/Qwen3.8-27B-4bit(affine 4 bit, group 64, 16.1 GB; baseQwen/Qwen3.8-27B). The Ollama NVFP4 weights do not load in mlx_lm (global scale). -
Server:
train/serve.sh(mlx_lm.server, port 8080, thinking on withreasoning_effortmedium, temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed. -
Eval subset (Kral decision):
train/subset.json(copyruns/stage1/subset.json), 25 tasks. The task document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after training. -
Runner:
train/baseline.py --label <label>(one task at a time, resultsruns/stage1/<label>.json, run directoriesruns/stage1/<label>/). -
Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule):
train/prepare.py, reporttrain/data/report.md. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens. Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed. Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82 pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834; no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %. Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003). 748 iterations for 2 epochs (the count is the same as with the first dedup rule by coincidence).test.jsonl= copy of valid.
Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; runs/archive/devstral/). Official baseline: runs/stage1/baseline.json (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in
pgrep -f mlx_lm), logruns/stage1/server.log. - Baseline on the 11-task subset (
train/subset.json), thinking off, loop guard 3, max_tokens 16384, started 2026-10-04 07:59 bytrain/baseline_chain.sh(run base 20500; the guard also ends read loops), logsruns/stage1/baseline_chain.logandruns/stage1/baseline.log. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended"). - Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (
_aborted_*inruns/stage1/baseline/, A4H objects deleted).
Next
-
After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
-
B2: read the last lines of
runs/stage1/baseline.log; summary fromruns/stage1/baseline.json(t01_test_budget40holds the T01 test result). Commit. -
Step 2, second part: valid loss of the base model (
mlx_lm.lora --testwithout adapter; check the options with--helpfirst). Add it toruns/stage1/baseline.json. -
Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first.
Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit) is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(
max_tokensonly for cloud models; the MLX server limit is--max-tokens 32768). - Session 2026-10-03 (evening): follow-ups of
docs/devir-notlari.mdsection 3 done from stored results (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait inruns/stage1/rerun_queue.txt(G0119, G0162) until the baseline ends.
Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No mlx_lm training.
- Base valid loss (Mac, MLX 4-bit, no adapter): 0.849, ppl 2.337 (
runs/stage1/baseline.json, keyvalid_loss). - Done: private dataset
erhankeseli/abap-stage1-data; private model repoerhankeseli/abap-stage1-adapter-test;train/hf_train.py(Unsloth job),train/peft_to_mlx.py(converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters. - Key names: PEFT
base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight(A: r x in) -> mlxlanguage_model.model.layers.N.<mod>.lora_a/b(transposed). mlx scale = alpha / r. - Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx
scale: 20would be alpha 320. - Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5): 304.5 s train = 30.5 s/step, job wall time 8 min 7 s = about 0.34 USD, peak GPU 45.4 GB. Adapter pushed to the private repo. Unsloth valid loss (bnb 4-bit): 0.920.
- Conversion test:
peft_to_mlx.pyconverted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors. Valid loss on the Mac with the adapter: 0.848 (base 0.849,runs/stage1/adapter_test_valid_loss.log). Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same. Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded). Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps). - Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length). Not started. Needs Kral's go and the alpha decision.
Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)
- Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18), alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
- Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
- Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30,
timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line
EVAL), checkpoints and valid loss every 200 steps inerhankeseli/abap-stage1-adapter-full/step<N>/(witheval.json). Expected about 6.3 h. - Rule: at step 200 convert the checkpoint (
train/peft_to_mlx.py), measure the valid loss on the Mac (runs/stage1/adapter_test_valid_loss.logcommand, about 20 min), compare the relative drop with the GPU value. Stop the job (hf jobs cancel) if they differ.
Step 200 check failed, full run cancelled (2026-10-04)
- GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
- Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log
runs/stage1/full_step200_valid_loss.log. - Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD).
Checkpoints
step200andstep400stay inerhankeseli/abap-stage1-adapter-full. - Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4 error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091 on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA, larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.
Findings and next GPU run (2026-10-05)
| Step | GPU valid loss (nf4 base) | GPU drop | Mac valid loss (MLX 4-bit, converted adapter) | Mac drop |
|---|---|---|---|---|
| 0 / base | 0.9228 | 0.849 | ||
| 200 | 0.6830 | -26.0 % | 0.758 | -10.7 % |
| 400 | 0.6003 | -35.0 % | 0.735 | -13.4 % |
- Step 400 Mac log:
runs/stage1/full_step400_valid_loss.log. The Mac drop grows with the steps (-10.7 % to -13.4 %), but stays far below the GPU drop. - Converter
train/peft_to_mlx.pyis proven (overfit test: Mac -98.5 %, GPU -99.97 %). The gap is the base mismatch: an nf4 adapter does not transfer well to the MLX 4-bit base. - Decisions (Kral, 2026-10-05): no more GPU runs now. Stage 1 is not trained alone. The stage 1 corpus is mixed with the stage 2 trajectories in one bf16 LoRA run later. The next GPU run uses a bf16 base (no nf4). A bf16 memory test is needed before it (27B bf16 weights are about 54 GB; check GPU size, sequence length 16384, gradient checkpointing).
- Tool names stay as in the EPOD ABAP MCP server (generic_v0 not used). The ADT "save failed" response cannot be fixed on the server side; the proxy adds a syntax check instead.
- HF cleanup done: test, overfit and full adapter repos deleted (step 200 and 400 checkpoints too). Kept: dataset
erhankeseli/abap-stage1-data. Local adapter folders anddata_overfitdeleted. BUDGET_LIMIT_USD= 161 (ledger 66.97 + 35 usage x 2.7), until 12 October.
Stage 2 data, Part B (2026-10-05)
-
Teacher budget:
BUDGET_LIMIT_USD161 (ledger). Until 11 October up to 35 usage (= 94.5 ledger) without asking. -
B1 generator training mode:
harness/trainset.py, pooltasks_gen/train(ids G1000+, run numbers 32000+), plan intasks_gen/train/plan.json(223 slots: category shares as the eval plan, 30 error tasks: named-type 10, reserved-word 10, long-names 10), overlap check against all eval tasks and earlier training tasks (harness/overlap.py: spec cosine 0.75, Goal + Business rules cosine 0.60, contract name Jaccard 0.60; calibrated on the eval pairs; K variants are the only eval pairs above 0.75). A too close bundle goes back to the model as a repair message. K variants are not made for training (they need the generic_v0 schema, which is not used). Started 2026-10-05 12:00 with 3 workers, target 200 accepted tasks, phase limit 27 ledger USD; logsruns/gen_train/w*.log, slot logstasks_gen/train/_logs/. -
B2a proxy: a write that fails with only "An error occured during the save operation" gets
syntaxCheckmessages from local abaplint parser (line + text). The EPOD syntax check cannot check the rejected source (tested 2026-10-05: it returned 0 errors for the stored stub), so abaplint is used. The raw result stays intrajectory.jsonl(raw_result). -
B2b record:
harness/record.pywritesrecord.json(task metadata, exact messages, tool schemas, raw tool results, metadata: score, cost, end reason, ...) andreasoning.json(teacher reasoning, not training data) per run. -
B2c converter:
train/to_qwen.py(Qwen 3.8 chat template,enable_thinking=False, Qwen XML tool call format; tokenizer round-trip check of every tool call; assistant spans for loss masking). B2d filter:train/accept.py. -
Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
-
B3 trajectory runner:
harness/trajectories.py(6 workers, DeepSeek V4.1 Flash, outputruns/traj/). Starts after generation. -
2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (
Runner.run(cds_calls=100),trajectories.CDS_CALLS). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted).metadata.tool_budgetis in every record. Failed CDS tasks get the second attempt with the new budget. -
2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = 2.07, not 2.7.
BUDGET_LIMIT_USD161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept).harness/ledger.pynow reads the limit from.envat each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs inruns/traj/_aborted/, rerun later).