Step 200 check: GPU -26.0 % vs Mac -10.7 %, full run cancelled at step ~418

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-04 20:48:49 +02:00
parent 0c0982f666
commit 39dcdd9846

View File

@@ -93,3 +93,15 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- Rule: at step 200 convert the checkpoint (`train/peft_to_mlx.py`), measure the valid loss on the Mac
(`runs/stage1/adapter_test_valid_loss.log` command, about 20 min), compare the relative drop with the GPU value.
Stop the job (`hf jobs cancel`) if they differ.
### Step 200 check failed, full run cancelled (2026-10-04)
- GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
- Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log `runs/stage1/full_step200_valid_loss.log`.
- Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD).
Checkpoints `step200` and `step400` stay in `erhankeseli/abap-stage1-adapter-full`.
- Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4
error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091
on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA,
larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.