Step 200 check: GPU -26.0 % vs Mac -10.7 %, full run cancelled at step ~418
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -93,3 +93,15 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
||||
- Rule: at step 200 convert the checkpoint (`train/peft_to_mlx.py`), measure the valid loss on the Mac
|
||||
(`runs/stage1/adapter_test_valid_loss.log` command, about 20 min), compare the relative drop with the GPU value.
|
||||
Stop the job (`hf jobs cancel`) if they differ.
|
||||
|
||||
### Step 200 check failed, full run cancelled (2026-10-04)
|
||||
|
||||
- GPU valid loss (Unsloth, bnb nf4 base): step 0 0.9228, step 200 0.6830 (-26.0 %), step 400 0.6003 (-35.0 %).
|
||||
- Mac (MLX affine 4-bit base) with the converted step-200 adapter: 0.849 -> 0.758 (-10.7 %), log `runs/stage1/full_step200_valid_loss.log`.
|
||||
- Rule (Opus): stop if the relative drops differ. Job 6ac28e19fbc85ba6823a0eef cancelled at step about 418 of 748 (about 4 h, roughly 10 USD).
|
||||
Checkpoints `step200` and `step400` stay in `erhankeseli/abap-stage1-adapter-full`.
|
||||
- Likely cause (not proven): the nf4 base is worse than the affine base (0.923 vs 0.849), and the adapter learns to repair the nf4
|
||||
error too. That part of the GPU drop does not exist on the Mac. Even so, the gain over the Mac base (0.849 -> 0.683 on GPU, 0.091
|
||||
on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
|
||||
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA,
|
||||
larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.
|
||||
|
||||
Reference in New Issue
Block a user