Overfit conversion test passed (Mac -98.5 %, GPU -99.97 %); full run started, alpha 32

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-04 19:34:24 +02:00
parent 20ef7b2331
commit 0c0982f666
2 changed files with 41 additions and 5 deletions

View File

@@ -80,3 +80,16 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps).
- Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length).
Not started. Needs Kral's go and the alpha decision.
## Overfit test and full run (2026-10-04, Opus 5.5 decision: alpha 32, cost limit 25 USD)
- Overfit test (job 6ac28c41404719ba3764f7fc): one train document (ZCL_DEMO_ABAP_STRUCTURES, 2503 tokens, train row 18),
alpha 32, lr 1e-3, 30 steps, no warmup, 5.8 s/step, about 0.4 USD. Unsloth loss on that document: 0.798 -> 0.000234 (-99.97 %).
- Mac with the converted adapter (scale 2.0): 0.811 (no adapter) -> 0.012 (-98.5 %). Conversion confirmed (a broken
conversion would stay near 0.8). The remaining gap is the base mismatch (bnb nf4 vs MLX affine 4-bit).
- Full run started: job 6ac28e19fbc85ba6823a0eef, a100-large, 748 steps, rank 16, alpha 32, lr 5e-5 cosine, warmup 30,
timeout 8 h (max 20 USD). Valid loss logged at step 0 (log line `EVAL`), checkpoints and valid loss every 200 steps in
`erhankeseli/abap-stage1-adapter-full/step<N>/` (with `eval.json`). Expected about 6.3 h.
- Rule: at step 200 convert the checkpoint (`train/peft_to_mlx.py`), measure the valid loss on the Mac
(`runs/stage1/adapter_test_valid_loss.log` command, about 20 min), compare the relative drop with the GPU value.
Stop the job (`hf jobs cancel`) if they differ.