180 lines
18 KiB
Markdown
180 lines
18 KiB
Markdown
# Stage 1 state
|
|
|
|
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:40.
|
|
|
|
## Done
|
|
|
|
- Step 0 (setup): `train/.venv` (Python 3.11.17, uv), `mlx-lm` 0.32.0 / `mlx` 0.32.3.
|
|
Model `mlx-community/Qwen3.8-27B-4bit` at `~/models/Qwen3.8-27B-4bit` (affine 4 bit, group 64, 16.1 GB;
|
|
base `Qwen/Qwen3.8-27B`). The Ollama NVFP4 weights do not load in mlx_lm (global scale).
|
|
- Server: `train/serve.sh` (`mlx_lm.server`, port 8080, thinking on with `reasoning_effort` medium,
|
|
temperature 0.2, top_p 0.95, top_k 20, min_p 0, max tokens 32768). Tool-call test passed.
|
|
- Eval subset (Kral decision): `train/subset.json` (copy `runs/stage1/subset.json`), 25 tasks. The task
|
|
document says "all accepted tasks" for step 2; Kral changed it to this subset. Use the same list after
|
|
training.
|
|
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
|
|
run directories `runs/stage1/<label>/`).
|
|
|
|
- Step 1 done (2026-10-03 22:20, Kral decisions applied, strict dedup rule): `train/prepare.py`, report
|
|
`train/data/report.md`. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
|
|
Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with
|
|
the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed.
|
|
Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82
|
|
pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834;
|
|
no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %.
|
|
Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
|
|
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
|
|
`test.jsonl` = copy of valid.
|
|
|
|
## Base model
|
|
|
|
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
|
|
|
|
## Running (detached) — historical, all ended
|
|
|
|
|
|
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
|
|
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
|
|
2026-10-04 07:59 by `train/baseline_chain.sh` (run base 20500; the guard also ends read loops), logs `runs/stage1/baseline_chain.log` and
|
|
`runs/stage1/baseline.log`. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended").
|
|
- Earlier attempts (T01 without the guard, T01+G0105 with the push-only guard) were stopped (`_aborted_*` in `runs/stage1/baseline/`, A4H objects deleted).
|
|
|
|
## Next
|
|
|
|
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
|
|
|
|
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
|
(`t01_test_budget40` holds the T01 test result). Commit.
|
|
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
|
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
|
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
|
iterations first.
|
|
|
|
## Notes
|
|
|
|
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 (MLX 4 bit)
|
|
is the new reference.
|
|
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
|
|
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
|
|
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
|
|
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
|
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
|
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|
|
|
|
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
|
|
|
|
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
|
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
|
|
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
|
|
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
|
|
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
|
|
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
|
|
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
|
|
- Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5):
|
|
304.5 s train = **30.5 s/step**, job wall time 8 min 7 s = about **0.34 USD**, peak GPU 45.4 GB. Adapter pushed to the private repo.
|
|
Unsloth valid loss (bnb 4-bit): 0.920.
|
|
- Conversion test: `peft_to_mlx.py` converted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors.
|
|
Valid loss on the Mac with the adapter: **0.848** (base 0.849, `runs/stage1/adapter_test_valid_loss.log`).
|
|
Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same.
|
|
Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded).
|
|
Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps).
|
|
- Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length).
|
|
Not started. Needs Kral's go and the alpha decision.
|
|
|
|
## Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)
|
|
|
|
- Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18),
|
|
alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
|
|
- Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken
|
|
conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
|
|
- Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30,
|
|
timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line `EVAL`), checkpoints and valid loss every 200 steps in
|
|
`erhankeseli/abap-stage1-adapter-full/step<N>/` (with `eval.json`). Expected about 6.3 h.
|
|
- Rule: at step 200 convert the checkpoint (`train/peft_to_mlx.py`), measure the valid loss on the Mac
|
|
(`runs/stage1/adapter_test_valid_loss.log` command, about 20 min), compare the relative drop with the GPU value.
|
|
Stop the job (`hf jobs cancel`) if they differ.
|
|
|
|
### Step 200 check failed, full run cancelled (2026-10-04)
|
|
|
|
- GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
|
|
- Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log `runs/stage1/full_step200_valid_loss.log`.
|
|
- Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD).
|
|
Checkpoints `step200` and `step400` stay in `erhankeseli/abap-stage1-adapter-full`.
|
|
- Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4
|
|
error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091
|
|
on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
|
|
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA,
|
|
larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.
|
|
|
|
## Findings and next GPU run (2026-10-05)
|
|
|
|
| Step | GPU valid loss (nf4 base) | GPU drop | Mac valid loss (MLX 4-bit, converted adapter) | Mac drop |
|
|
|---|---|---|---|---|
|
|
| 0 / base | 0.9228 | | 0.849 | |
|
|
| 200 | 0.6830 | -26.0 % | 0.758 | -10.7 % |
|
|
| 400 | 0.6003 | -35.0 % | 0.735 | -13.4 % |
|
|
|
|
- Step 400 Mac log: `runs/stage1/full_step400_valid_loss.log`. The Mac drop grows with the steps (-10.7 % to -13.4 %), but stays far below the GPU drop.
|
|
- Converter `train/peft_to_mlx.py` is proven (overfit test: Mac -98.5 %, GPU -99.97 %). The gap is the base mismatch: an nf4 adapter does not transfer well to the MLX 4-bit base.
|
|
- Decisions (Kral, 2026-10-05): no more GPU runs now. Stage 1 is not trained alone. The stage 1 corpus is mixed with the stage 2 trajectories in one bf16 LoRA run later. The next GPU run uses a bf16 base (no nf4). A bf16 memory test is needed before it (27B bf16 weights are about 54 GB; check GPU size, sequence length 16384, gradient checkpointing).
|
|
- Tool names stay as in the EPOD ABAP MCP server (generic_v0 not used). The ADT "save failed" response cannot be fixed on the server side; the proxy adds a syntax check instead.
|
|
- HF cleanup done: test, overfit and full adapter repos deleted (step 200 and 400 checkpoints too). Kept: dataset `erhankeseli/abap-stage1-data`. Local adapter folders and `data_overfit` deleted.
|
|
- `BUDGET_LIMIT_USD` = 161 (ledger 66.97 + 35 usage x 2.7), until 12 October.
|
|
|
|
## Stage 2 data, Part B (2026-10-05)
|
|
|
|
- Teacher budget: `BUDGET_LIMIT_USD` 161 (ledger). Until 11 October up to 35 usage (= 94.5 ledger) without asking.
|
|
- B1 generator training mode: `harness/trainset.py`, pool `tasks_gen/train` (ids G1000+, run numbers 32000+), plan in
|
|
`tasks_gen/train/plan.json` (223 slots: category shares as the eval plan, 30 error tasks: named-type 10, reserved-word 10,
|
|
long-names 10), overlap check against all eval tasks and earlier training tasks (`harness/overlap.py`: spec cosine 0.75,
|
|
Goal + Business rules cosine 0.60, contract name Jaccard 0.60; calibrated on the eval pairs; K variants are the only
|
|
eval pairs above 0.75). A too close bundle goes back to the model as a repair message. K variants are not made for
|
|
training (they need the generic_v0 schema, which is not used). Started 2026-10-05 12:00 with 3 workers, target 200
|
|
accepted tasks, phase limit 27 ledger USD; logs `runs/gen_train/w*.log`, slot logs `tasks_gen/train/_logs/`.
|
|
- B2a proxy: a write that fails with only "An error occured during the save operation" gets `syntaxCheck` messages from local
|
|
abaplint parser (line + text). The EPOD syntax check cannot check the rejected source (tested 2026-10-05: it returned
|
|
0 errors for the stored stub), so abaplint is used. The raw result stays in `trajectory.jsonl` (`raw_result`).
|
|
- B2b record: `harness/record.py` writes `record.json` (task metadata, exact messages, tool schemas, raw tool results,
|
|
metadata: score, cost, end reason, ...) and `reasoning.json` (teacher reasoning, not training data) per run.
|
|
- B2c converter: `train/to_qwen.py` (Qwen 3.8 chat template, `enable_thinking=False`, Qwen XML tool call format; tokenizer
|
|
round-trip check of every tool call; assistant spans for loss masking). B2d filter: `train/accept.py`.
|
|
- Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
|
|
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
|
|
|
|
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.
|
|
|
|
- 2026-10-05 14:50 budget correction: Ollama panel 30.98 of 60 usage at ledger 80.1; since 24.63 (ledger 66.97) the ratio is 13.1 ledger / 6.35 usage = **2.07**, not 2.7. `BUDGET_LIMIT_USD` 161 -> 142 (guard at 134 ledger = panel about 57 usage, reserve 8 ledger kept). `harness/ledger.py` now reads the limit from `.env` at each check. Pipeline restarted (2 interrupted runs cleaned from A4H, dirs in `runs/traj/_aborted/`, rerun later).
|
|
|
|
- 2026-10-05 15:55 budget correction 2: panel 34.79 at ledger 85.45. Ledger per usage is not constant: 2.04 (24.63 to 31.48, mostly generation) and 1.35 (31.48 to 34.79, mostly trajectory runs). Trajectory runs cost more real usage per ledger USD. Guard set from the panel: panel left 60 - 34.79 - 3.0 reserve = 22.2 usage x 1.35 = 30 ledger, so guard 115 ledger, `BUDGET_LIMIT_USD` 123 (reserve 8). Rule: Kral gives the panel value now and then; the limit is recomputed from it.
|
|
|
|
- 2026-10-05 17:50 budget correction 3: panel 41.07 at ledger 93.45; since 34.79 (85.45): 8.0 ledger / 6.28 usage = 1.27 (trajectory phase). Guard recomputed with ratio 1.2: panel left 60 - 41.07 - 3.0 reserve = 15.9 usage x 1.2 = 19 ledger, guard 112.6, `BUDGET_LIMIT_USD` 121. Burn about 3.7 usage per hour: the guard is reached around 22:00 on 2026-10-05.
|
|
|
|
### Stage 2 summary at 50 accepted trajectories (2026-10-05 18:02)
|
|
|
|
- Tasks: 139 accepted of 166 generated (K variants 18); trajectory runs 59, accepted trajectories 50 (acceptance 85%).
|
|
- Budget: used 10.2 usage (ledger 27.5) since 2026-10-05; left to the guard about 15 usage (ledger 18.5 to guard 113; reserve 8 ledger kept; corrected 2026-10-05, the running controller read the old limit 142); allowance until 11 October 24.8 usage.
|
|
- Repair share (accepted trajectories with an error followed by a fix): 32/50 = 64%.
|
|
- By category, accepted/runs: A 7/7 (repair 4); B 5/8 (repair 5); C 7/8 (repair 6); D 6/6 (repair 3); E 9/10 (repair 6); F 6/6 (repair 2); G 6/6 (repair 5); H 3/5 (repair 0); I 1/3 (repair 1).
|
|
- By object type, accepted/runs: CLAS 38/42 (repair 21); DDLS 6/11 (repair 6); FUNC 6/6 (repair 5).
|
|
- Tokens of accepted samples (20 tool schemas kept): p50 20147, p90 39458, p95 47694, max 71791, n 50.
|
|
- Syntax hints (proxy syntaxCheck added): 1 in 1 runs.
|
|
- Harness events: 2 ({'text:[LOCK]': 1, 'teardown': 1}); trajectory workers now 1.
|
|
|
|
### Notes after the Opus review of the first stage 2 summary (2026-10-05 18:30)
|
|
|
|
- **Token length (for the bf16 memory test):** accepted samples, 20 tool schemas kept: p50 20k, p90 39k, **p95 48k, max 72k** (about 10k tokens are the tool schemas). The bf16 memory test must use a **sequence length of 48k** (covers 95 % of the samples; samples over 48k are cut or dropped; 72k is not needed in the test).
|
|
- **DDLS reject reasons** (11 CDS runs: 6 accepted, 5 rejected): tool budget 2 (old limit 60, before the 100 budget), empty response 3 (one turn reached the 32k output limit while the model wrote its own CDS test class). Activation 0, hidden tests 0, ATC 0: all 11 runs passed every gate and every hidden test. Failed CDS tasks get the second attempt (pipeline rule).
|
|
- **G1034 lock event (cause not found):** after `sap_push_source includeType=testclasses` the class stayed locked ("User KESELI is currently editing", SM12 entry needed; teardown said "You are already editing"). Not a harness bug: the lock outlived the MCP session and the run. Not reproduced in 3 tests on A4H (alone, 60 writes overlapping 2690 unit test calls of another session, writes with ATC + coverage + member listing + a second session; `push_element` on a test-class method). One case in about 90 runs. Details for the EPOD developer: `docs/epod-lock-leak.md`. The harness treats it as an infrastructure event (run not accepted, worker count 2 to 1, 3 equal events stop the pipeline) and lists the object in the dashboard.
|
|
|
|
### Object type mix fix (2026-10-05 19:30, Opus review item 1)
|
|
|
|
- **Cause:** plans 1 and 2 were built from `evalset.SLOTS`, which has only CLAS, FUNC, PROG and DDLS. The mix of CLAUDE.md section 1 (INTF, DDIC, MSAG, exception) was never in a plan, so no slot existed for those types and nothing was rejected. The harness also lacked G2 checks, message writing and mutants for them (`evalset.py` header said so).
|
|
- **Before (139 accepted tasks):** CLAS 49 %, FUNC 21 %, DDLS 19 %, PROG 9 %, INTF/TABL/STRU/MSAG 0, exception 2 %. Target: CLAS 28, INTF 7, DDLS 25, FUNC 15, PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5 (percent; `harness/mix.py`).
|
|
- **Fixes:** G2 for TABL/STRU fields and MSAG message numbers, `sap_push_message` in the oracle and setup, mutants for DDIC, interfaces, message classes and exception classes, G6 cascade fix (abaplint does not know CX_ super classes), format notes per type in the generator, `implements` may be a list. DTEL/DOMA are not made (DDIC share is TABL 8 + STRU 2).
|
|
- **First tasks per new type, all accepted with oracle 100, null 0 and mutation kill rate 100 %:** G1900 INTF (3 tries), G1901 TABL (1), G1902 MSAG (2), G1903 STRU (3), G1904 exception (mutation rerun after the exception mutants). Model mistakes fed back by the generator: unit/currency reference annotation in DDL needs `'table.field'`, RTTI length is in bytes.
|
|
- **Generation:** `trainset run --plan balanced` (3 workers, started by the pipeline): every slot takes the kind with the biggest deficit against the target; a kind is skipped with 6+ tries and under 20 % accepted, or with more than 8 accepted tasks waiting for a first trajectory run. 20 % error-targeted, 30 % hard (difficulty 3). K variants stay at 10 %.
|
|
- **Trajectories:** the next task is the one whose kind has the biggest deficit (accepted trajectories per kind). Summary now lists the mix and DDLS reject reasons.
|
|
- **Incident 19:25:** restarting the controller I started a second one by mistake (wrong `pgrep` pattern) and deleted the objects of a running run. Both controllers were stopped, leftovers cleaned, one controller runs. One table `Z4AJ50UB_PO_HEAD` kept a lock from the interrupted write (SM12 needed). Interrupting a write leaks the lock: never kill a controller during a run without checking.
|
|
|
|
- 2026-10-05 21:31 budget correction 4: panel 50.00 at ledger 111.22; since 45.0 (100.62): 10.6 ledger / 5.0 usage = 2.12 (generation of new types and first runs). Guard recomputed with the pessimistic ratio 1.2: panel left 60 - 50 - 3.0 reserve = 7.0 usage x 1.2 = 8.4 ledger, guard 119.6, `BUDGET_LIMIT_USD` 128 (reserve 8). Only kinds below target run (Kral + Opus 2026-10-05).
|