diff --git a/train/STATE.md b/train/STATE.md index 64d5c43..2ed9c24 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -70,5 +70,13 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training. - Key names: PEFT `base_model.model.model.language_model.layers.N..lora_A/B.weight` (A: r x in) -> mlx `language_model.model.layers.N..lora_a/b` (transposed). mlx scale = alpha / r. - Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320. -- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started. -- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost. +- Pipeline test done (2026-10-04, job 6ac2845ffbc85ba6823a0856, a100-large, 10 steps, rank 16, alpha 16, lr 5e-5): + 304.5 s train = **30.5 s/step**, job wall time 8 min 7 s = about **0.34 USD**, peak GPU 45.4 GB. Adapter pushed to the private repo. + Unsloth valid loss (bnb 4-bit): 0.920. +- Conversion test: `peft_to_mlx.py` converted 800 tensors (64 layers, 10 module types); mlx-lm loaded them without shape errors. + Valid loss on the Mac with the adapter: **0.848** (base 0.849, `runs/stage1/adapter_test_valid_loss.log`). + Weak test: 10 steps move the loss by 0.001 only, so a wrong conversion with a small effect looks the same. + Unsloth 0.920 and Mac 0.848 are not comparable (bnb nf4 vs MLX affine 4-bit base; no Unsloth base loss recorded). + Stronger check for later: log the Unsloth step-0 loss, or compare the Mac loss after a longer run (the first 200 steps). +- Full run estimate: 748 steps x 30.5 s = about 6.3 h = about 16 USD (range 12-21 USD, step time varies with document length). + Not started. Needs Kral's go and the alpha decision.