Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -26,7 +26,12 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
|
||||
`test.jsonl` = copy of valid.
|
||||
|
||||
## Running (detached)
|
||||
## Base model
|
||||
|
||||
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
|
||||
|
||||
## Running (detached) — historical, all ended
|
||||
|
||||
|
||||
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
|
||||
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
|
||||
@@ -40,8 +45,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
|
||||
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
||||
(`t01_test_budget40` holds the T01 test result). Commit.
|
||||
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
|
||||
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
||||
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
||||
iterations first.
|
||||
@@ -56,3 +60,15 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
||||
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
||||
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|
||||
|
||||
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
|
||||
|
||||
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
||||
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
|
||||
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
|
||||
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
|
||||
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
|
||||
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
|
||||
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
|
||||
- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started.
|
||||
- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost.
|
||||
|
||||
Reference in New Issue
Block a user