20 KiB
Stage 1 state
Task: docs/stage1-training-task.md. Settings and weights: train/README.md. Updated 2026-10-03 22:40.
Done
-
Step 0 (setup):
train/.venv(Python 3.11.17, uv),mlx-lm0.32.0 /mlx0.32.3. Modelmlx-community/Qwen3.8-27B-4bitat~/models/Qwen3.8-27B-4bit(affine 4 bit, group 64, 16.1 GB; baseQwen/Qwen3.8-27B). The Ollama NVFP4 weights do not load in mlx_lm (global scale). -
Server:
train/serve.sh(mlx_lm.server, port 8080, thinking on withreasoning_effortmedium, temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed. -
Eval subset (Kral decision):
train/subset.json(copyruns/stage1/subset.json), 25 tasks. The task document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after training. -
Runner:
train/baseline.py --label <label>(one task at a time, resultsruns/stage1/<label>.json, run directoriesruns/stage1/<label>/). -
Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule):
train/prepare.py, reporttrain/data/report.md. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens. Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed. Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82 pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834; no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %. Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003). 748 iterations for 2 epochs (the count is the same as with the first dedup rule by coincidence).test.jsonl= copy of valid.
Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; runs/archive/devstral/). Official baseline: runs/stage1/baseline.json (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in
pgrep -f mlx_lm), logruns/stage1/server.log. - Baseline on the 11-task subset (
train/subset.json), thinking off, loop guard 3, max_tokens 16384, started 2026-10-04 07:59 bytrain/baseline_chain.sh(run base 20500; the guard also ends read loops), logsruns/stage1/baseline_chain.logandruns/stage1/baseline.log. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended"). - Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (
_aborted_*inruns/stage1/baseline/, A4H objects deleted).
Next
-
After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
-
B2: read the last lines of
runs/stage1/baseline.log; summary fromruns/stage1/baseline.json(t01_test_budget40holds the T01 test result). Commit. -
Step 2, second part: valid loss of the base model (
mlx_lm.lora --testwithout adapter; check the options with--helpfirst). Add it toruns/stage1/baseline.json. -
Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first.
Notes
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit) is the new reference.
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
(
max_tokensonly for cloud models; the MLX server limit is--max-tokens 32768). - Session 2026-10-03 (evening): follow-ups of
docs/devir-notlari.mdsection 3 done from stored results (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait inruns/stage1/rerun_queue.txt(G0119, G0162) until the baseline ends.
Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No mlx_lm training.
- Base valid loss (Mac, MLX 4-bit, no adapter): 0.849, ppl 2.337 (
runs/stage1/baseline.json, keyvalid_loss). - Done: private dataset
erhankeseli/abap-stage1-data; private model repoerhankeseli/abap-stage1-adapter-test;train/hf_train.py(Unsloth job),train/peft_to_mlx.py(converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters. - Key names: PEFT
base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight(A: r x in) -> mlxlanguage_model.model.layers.N.<mod>.lora_a/b(transposed). mlx scale = alpha / r. - Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx
scale: 20would be alpha 320. - Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5): 304.5 s train = 30.5 s/step, job wall time 8 min 7 s = about 0.34 USD, peak GPU 45.4 GB. Adapter pushed to the private repo. Unsloth valid loss (bnb 4-bit): 0.920.
- Conversion test:
peft_to_mlx.pyconverted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors. Valid loss on the Mac with the adapter: 0.848 (base 0.849,runs/stage1/adapter_test_valid_loss.log). Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same. Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded). Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps). - Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length). Not started. Needs Kral's go and the alpha decision.
Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)
- Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18), alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
- Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
- Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30,
timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line
EVAL), checkpoints and valid loss every 200 steps inerhankeseli/abap-stage1-adapter-full/step<N>/(witheval.json). Expected about 6.3 h. - Rule: at step 200 convert the checkpoint (
train/peft_to_mlx.py), measure the valid loss on the Mac (runs/stage1/adapter_test_valid_loss.logcommand, about 20 min), compare the relative drop with the GPU value. Stop the job (hf jobs cancel) if they differ.
Step 200 check failed, full run cancelled (2026-10-04)
- GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
- Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log
runs/stage1/full_step200_valid_loss.log. - Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD).
Checkpoints
step200andstep400stay inerhankeseli/abap-stage1-adapter-full. - Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4 error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091 on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA, larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.
Findings and next GPU run (2026-10-05)
| Step | GPU valid loss (nf4 base) | GPU drop | Mac valid loss (MLX 4-bit, converted adapter) | Mac drop |
|---|---|---|---|---|
| 0 / base | 0.9228 | 0.849 | ||
| 200 | 0.6830 | -26.0 % | 0.758 | -10.7 % |
| 400 | 0.6003 | -35.0 % | 0.735 | -13.4 % |
- Step 400 Mac log:
runs/stage1/full_step400_valid_loss.log. The Mac drop grows with the steps (-10.7 % to -13.4 %), but stays far below the GPU drop. - Converter
train/peft_to_mlx.pyis proven (overfit test: Mac -98.5 %, GPU -99.97 %). The gap is the base mismatch: an nf4 adapter does not transfer well to the MLX 4-bit base. - Decisions (Kral, 2026-10-05): no more GPU runs now. Stage 1 is not trained alone. The stage 1 corpus is mixed with the stage 2 trajectories in one bf16 LoRA run later. The next GPU run uses a bf16 base (no nf4). A bf16 memory test is needed before it (27B bf16 weights are about 54 GB; check GPU size, sequence length 16384, gradient checkpointing).
- Tool names stay as in the EPOD ABAP MCP server (generic_v0 not used). The ADT "save failed" response cannot be fixed on the server side; the proxy adds a syntax check instead.
- HF cleanup done: test, overfit and full adapter repos deleted (step 200 and 400 checkpoints too). Kept: dataset
erhankeseli/abap-stage1-data. Local adapter folders anddata_overfitdeleted. BUDGET_LIMIT_USD= 161 (ledger 66.97 + 35 usage x 2.7), until 12 October.
Stage 2 data, Part B (2026-10-05)
-
Teacher budget:
BUDGET_LIMIT_USD161 (ledger). Until 11 October up to 35 usage (= 94.5 ledger) without asking. -
B1 generator training mode:
harness/trainset.py, pooltasks_gen/train(ids G1000+, run numbers 32000+), plan intasks_gen/train/plan.json(223 slots: category shares as the eval plan, 30 error tasks: named-type 10, reserved-word 10, long-names 10), overlap check against all eval tasks and earlier training tasks (harness/overlap.py: spec cosine 0.75, Goal + Business rules cosine 0.60, contract name Jaccard 0.60; calibrated on the eval pairs; K variants are the only eval pairs above 0.75). A too close bundle goes back to the model as a repair message. K variants are not made for training (they need the generic_v0 schema, which is not used). Started 2026-10-05 12:00 with 3 workers, target 200 accepted tasks, phase limit 27 ledger USD; logsruns/gen_train/w*.log, slot logstasks_gen/train/_logs/. -
B2a proxy: a write that fails with only "An error occured during the save operation" gets
syntaxCheckmessages from local abaplint parser (line + text). The EPOD syntax check cannot check the rejected source (tested 2026-10-05: it returned 0 errors for the stored stub), so abaplint is used. The raw result stays intrajectory.jsonl(raw_result). -
B2b record:
harness/record.pywritesrecord.json(task metadata, exact messages, tool schemas, raw tool results, metadata: score, cost, end reason, ...) andreasoning.json(teacher reasoning, not training data) per run. -
B2c converter:
train/to_qwen.py(Qwen 3.8 chat template,enable_thinking=False, Qwen XML tool call format; tokenizer round-trip check of every tool call; assistant spans for loss masking). B2d filter:train/accept.py. -
Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
-
B3 trajectory runner:
harness/trajectories.py(6 workers, DeepSeek V4.1 Flash, outputruns/traj/). Starts after generation. -
2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (
Runner.run(cds_calls=100),trajectories.CDS_CALLS). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted).metadata.tool_budgetis in every record. Failed CDS tasks get the second attempt with the new budget. -
2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = 2.07, not 2.7.
BUDGET_LIMIT_USD161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept).harness/ledger.pynow reads the limit from.envat each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs inruns/traj/_aborted/, rerun later). -
2026-10-05 15:55 budget correction 2: panel 34.79 at ledger 85.45. Ledger per usage is not constant: 2.04 (24.63 to 31.48, mostly generation) and 1.35 (31.48 to 34.79, mostly trajectory runs). Trajectory runs cost more real usage per ledger USD. Guard set from the panel: panel left 60 - 34.79 - 3.0 reserve = 22.2 usage x 1.35 = 30 ledger, so guard 115 ledger,
BUDGET_LIMIT_USD123 (reserve 8). Rule: Kral gives the panel value now and then; the limit is recomputed from it. -
2026-10-05 17:50 budget correction 3: panel 41.07 at ledger 93.45; since 34.79 (85.45): 8.0 ledger / 6.28 usage = 1.27 (trajectory phase). Guard recomputed with ratio 1.2: panel left 60 - 41.07 - 3.0 reserve = 15.9 usage x 1.2 = 19 ledger, guard 112.6,
BUDGET_LIMIT_USD121. Burn about 3.7 usage per hour: the guard is reached around 22:00 on 2026-10-05.
Stage 2 summary at 50 accepted trajectories (2026-10-05 18:02)
- Tasks: 139 accepted of 166 generated (K variants 18); trajectory runs 59, accepted trajectories 50 (acceptance 85%).
- Budget: used 10.2 usage (ledger 27.5) since 2026-10-05; left to the guard about 15 usage (ledger 18.5 to guard 113; reserve 8 ledger kept; corrected 2026-10-05, the running controller read the old limit 142); allowance until 11 October 24.8 usage.
- Repair share (accepted trajectories with an error followed by a fix): 32/50 = 64%.
- By category, accepted/runs: A 7/7 (repair 4); B 5/8 (repair 5); C 7/8 (repair 6); D 6/6 (repair 3); E 9/10 (repair 6); F 6/6 (repair 2); G 6/6 (repair 5); H 3/5 (repair 0); I 1/3 (repair 1).
- By object type, accepted/runs: CLAS 38/42 (repair 21); DDLS 6/11 (repair 6); FUNC 6/6 (repair 5).
- Tokens of accepted samples (20 tool schemas kept): p50 20147, p90 39458, p95 47694, max 71791, n 50.
- Syntax hints (proxy syntaxCheck added): 1 in 1 runs.
- Harness events: 2 ({'text:[LOCK]': 1, 'teardown': 1}); trajectory workers now 1.
Notes after the Opus review of the first stage 2 summary (2026-10-05 18:30)
- Token length (for the bf16 memory test): accepted samples, 20 tool schemas kept: p50 20k, p90 39k, p95 48k, max 72k (about 10k tokens are the tool schemas). The bf16 memory test must use a sequence length of 48k (covers 95 % of the samples; samples over 48k are cut or dropped; 72k is not needed in the test).
- DDLS reject reasons (11 CDS runs: 6 accepted, 5 rejected): tool budget 2 (old limit 60, before the 100 budget), empty response 3 (one turn reached the 32k output limit while the model wrote its own CDS test class). Activation 0, hidden tests 0, ATC 0: all 11 runs passed every gate and every hidden test. Failed CDS tasks get the second attempt (pipeline rule).
- G1034 lock event (cause not found): after
sap_push_source includeType=testclassesthe class stayed locked ("User KESELI is currently editing", SM12 entry needed; teardown said "You are already editing"). Not a harness bug: the lock outlived the MCP session and the run. Not reproduced in 3 tests on A4H (alone, 60 writes overlapping 2690 unit test calls of another session, writes with ATC + coverage + member listing + a second session;push_elementon a test-class method). One case in about 90 runs. Details for the EPOD developer:docs/epod-lock-leak.md. The harness treats it as an infrastructure event (run not accepted, worker count 2 to 1, 3 equal events stop the pipeline) and lists the object in the dashboard.
Object type mix fix (2026-10-05 19:30, Opus review item 1)
-
Cause: plans 1 and 2 were built from
evalset.SLOTS, which has only CLAS, FUNC, PROG and DDLS. The mix of CLAUDE.md section 1 (INTF, DDIC, MSAG, exception) was never in a plan, so no slot existed for those types and nothing was rejected. The harness also lacked G2 checks, message writing and mutants for them (evalset.pyheader said so). -
Before (139 accepted tasks): CLAS 49 %, FUNC 21 %, DDLS 19 %, PROG 9 %, INTF/TABL/STRU/MSAG 0, exception 2 %. Target: CLAS 28, INTF 7, DDLS 25, FUNC 15, PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5 (percent;
harness/mix.py). -
Fixes: G2 for TABL/STRU fields and MSAG message numbers,
sap_push_messagein the oracle and setup, mutants for DDIC, interfaces, message classes and exception classes, G6 cascade fix (abaplint does not know CX_ super classes), format notes per type in the generator,implementsmay be a list. DTEL/DOMA are not made (DDIC share is TABL 8 + STRU 2). -
First tasks per new type, all accepted with oracle 100, null 0 and mutation kill rate 100 %: G1900 INTF (3 tries), G1901 TABL (1), G1902 MSAG (2), G1903 STRU (3), G1904 exception (mutation rerun after the exception mutants). Model mistakes fed back by the generator: unit/currency reference annotation in DDL needs
'table.field', RTTI length is in bytes. -
Generation:
trainset run --plan balanced(3 workers, started by the pipeline): every slot takes the kind with the biggest deficit against the target; a kind is skipped with 6+ tries and under 20 % accepted, or with more than 8 accepted tasks waiting for a first trajectory run. 20 % error-targeted, 30 % hard (difficulty 3). K variants stay at 10 %. -
Trajectories: the next task is the one whose kind has the biggest deficit (accepted trajectories per kind). Summary now lists the mix and DDLS reject reasons.
-
Incident 19:25: restarting the controller I started a second one by mistake (wrong
pgreppattern) and deleted the objects of a running run. Both controllers were stopped, leftovers cleaned, one controller runs. One tableZ4AJ50UB_PO_HEADkept a lock from the interrupted write (SM12 needed). Interrupting a write leaks the lock: never kill a controller during a run without checking. -
2026-10-05 21:31 budget correction 4: panel 50.00 at ledger 111.22; since 45.0 (100.62): 10.6 ledger / 5.0 usage = 2.12 (generation of new types and first runs). Guard recomputed with the pessimistic ratio 1.2: panel left 60 - 50 - 3.0 reserve = 7.0 usage x 1.2 = 8.4 ledger, guard 119.6,
BUDGET_LIMIT_USD128 (reserve 8). Only kinds below target run (Kral + Opus 2026-10-05). -
2026-10-05 22:35 pipeline stopped at 21:43 by an infrastructure outage, found 22:26 (my miss). The MCP server (127.0.0.1:3000) refused connections for a short time; three runs got
URLError(ConnectionRefusedError)and the pipeline counted three equal harness errors and stopped (rule). Not a model or harness bug: A4H was up (11 h), the MCP server answers again. Newharness/infra.py: an outage of MCP or A4H is waited out (check every 30 s, two answers in a row), the run's objects are deleted, the same run starts again; it is not counted as an event or a failure (runs/traj/outages.jsonl). Same for series A (applies after a restart ofharness.localqwen; the running series keeps the old code and would skip the task). The 3 interrupted runs were cleaned (11 objects) and get their second attempt. Budget: panel 50.84 at ledger 113.36 (2.54 ledger per usage in the last stretch); guard recomputed: panel left 9.16 - 3.0 reserve = 6.2 usage x 1.2 = 7.4 ledger,BUDGET_LIMIT_USD129 (guard 121). Dashboard: a stopped pipeline shows no burn rate or ETA. -
2026-10-05 22:40 cause of the 21:43 outage confirmed by Kral: he restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment; their objects could be deleted afterwards (no stale lock), so a restart of the MCP server in the middle of a write did not leak a lock in this case. Rule: restarting Eclipse is fine now (the pipeline waits and reruns), but check
python3 scripts_probe/lockprobe.pyfor stale locks afterwards.