Files
abap-llm/train/STATE.md

145 lines
12 KiB
Markdown

# Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:40.
## Done
- Step 0 (setup): `train/.venv` (Python 3.11.17, uv), `mlx-lm` 0.32.0 / `mlx` 0.32.3.
Model `mlx-community/Qwen3.8-27B-4bit` at `~/models/Qwen3.8-27B-4bit` (affine 4 bit, group 64, 16.1 GB;
base `Qwen/Qwen3.8-27B`). The Ollama NVFP4 weights do not load in mlx_lm (global scale).
- Server: `train/serve.sh` (`mlx_lm.server`, port 8080, thinking on with `reasoning_effort` medium,
temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.
- Eval subset (Kral decision): `train/subset.json` (copy `runs/stage1/subset.json`), 25 tasks. The task
document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after
training.
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule): `train/prepare.py`, report
`train/data/report.md`. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with
the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed.
Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82
pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834;
no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %.
Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
`test.jsonl` = copy of valid.
## Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
## Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
2026-10-04 07:59 by `train/baseline_chain.sh` (run base 20500; the guard also ends read loops), logs `runs/stage1/baseline_chain.log` and
`runs/stage1/baseline.log`. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended").
- Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (`_aborted_*` in `runs/stage1/baseline/`, A4H objects deleted).
## Next
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
iterations first.
## Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit)
is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
- Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5):
304.5 s train = **30.5 s/step**, job wall time 8 min 7 s = about **0.34 USD**, peak GPU 45.4 GB. Adapter pushed to the private repo.
Unsloth valid loss (bnb 4-bit): 0.920.
- Conversion test: `peft_to_mlx.py` converted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors.
Valid loss on the Mac with the adapter: **0.848** (base 0.849, `runs/stage1/adapter_test_valid_loss.log`).
Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same.
Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded).
Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps).
- Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length).
Not started. Needs Kral's go and the alpha decision.
## Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)
- Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18),
alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
- Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken
conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
- Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30,
timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line `EVAL`), checkpoints and valid loss every 200 steps in
`erhankeseli/abap-stage1-adapter-full/step<N>/` (with `eval.json`). Expected about 6.3 h.
- Rule: at step 200 convert the checkpoint (`train/peft_to_mlx.py`), measure the valid loss on the Mac
(`runs/stage1/adapter_test_valid_loss.log` command, about 20 min), compare the relative drop with the GPU value.
Stop the job (`hf jobs cancel`) if they differ.
### Step 200 check failed, full run cancelled (2026-10-04)
- GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
- Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log `runs/stage1/full_step200_valid_loss.log`.
- Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD).
Checkpoints `step200` and `step400` stay in `erhankeseli/abap-stage1-adapter-full`.
- Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4
error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091
on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA,
larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.
## Findings and next GPU run (2026-10-05)
| Step | GPU valid loss (nf4 base) | GPU drop | Mac valid loss (MLX 4-bit, converted adapter) | Mac drop |
|---|---|---|---|---|
| 0 / base | 0.9228 | | 0.849 | |
| 200 | 0.6830 | -26.0 % | 0.758 | -10.7 % |
| 400 | 0.6003 | -35.0 % | 0.735 | -13.4 % |
- Step 400 Mac log: `runs/stage1/full_step400_valid_loss.log`. The Mac drop grows with the steps (-10.7 % to -13.4 %), but stays far below the GPU drop.
- Converter `train/peft_to_mlx.py` is proven (overfit test: Mac -98.5 %, GPU -99.97 %). The gap is the base mismatch: an nf4 adapter does not transfer well to the MLX 4-bit base.
- Decisions (Kral, 2026-10-05): no more GPU runs now. Stage 1 is not trained alone. The stage 1 corpus is mixed with the stage 2 trajectories in one bf16 LoRA run later. The next GPU run uses a bf16 base (no nf4). A bf16 memory test is needed before it (27B bf16 weights are about 54 GB; check GPU size, sequence length 16384, gradient checkpointing).
- Tool names stay as in the EPOD ABAP MCP server (generic_v0 not used). The ADT "save failed" response cannot be fixed on the server side; the proxy adds a syntax check instead.
- HF cleanup done: test, overfit and full adapter repos deleted (step 200 and 400 checkpoints too). Kept: dataset `erhankeseli/abap-stage1-data`. Local adapter folders and `data_overfit` deleted.
- `BUDGET_LIMIT_USD` = 161 (ledger 66.97 + 35 usage x 2.7), until 12 October.
## Stage 2 data, Part B (2026-10-05)
- Teacher budget: `BUDGET_LIMIT_USD` 161 (ledger). Until 11 October up to 35 usage (= 94.5 ledger) without asking.
- B1 generator training mode: `harness/trainset.py`, pool `tasks_gen/train` (ids G1000+, run numbers 32000+), plan in
`tasks_gen/train/plan.json` (223 slots: category shares as the eval plan, 30 error tasks: named-type 10, reserved-word 10,
long-names 10), overlap check against all eval tasks and earlier training tasks (`harness/overlap.py`: spec cosine 0.75,
Goal + Business rules cosine 0.60, contract name Jaccard 0.60; calibrated on the eval pairs; K variants are the only
eval pairs above 0.75). A too close bundle goes back to the model as a repair message. K variants are not made for
training (they need the generic_v0 schema, which is not used). Started 2026-10-05 12:00 with 3 workers, target 200
accepted tasks, phase limit 27 ledger USD; logs `runs/gen_train/w*.log`, slot logs `tasks_gen/train/_logs/`.
- B2a proxy: a write that fails with only "An error occured during the save operation" gets `syntaxCheck` messages from local
abaplint parser (line + text). The EPOD syntax check cannot check the rejected source (tested 2026-10-05: it returned
0 errors for the stored stub), so abaplint is used. The raw result stays in `trajectory.jsonl` (`raw_result`).
- B2b record: `harness/record.py` writes `record.json` (task metadata, exact messages, tool schemas, raw tool results,
metadata: score, cost, end reason, ...) and `reasoning.json` (teacher reasoning, not training data) per run.
- B2c converter: `train/to_qwen.py` (Qwen 3.8 chat template, `enable_thinking=False`, Qwen XML tool call format; tokenizer
round-trip check of every tool call; assistant spans for loss masking). B2d filter: `train/accept.py`.
- Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.