Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-04 18:51:40 +02:00
parent a2ba9e7b44
commit 5447874fd3
8 changed files with 196 additions and 29 deletions

View File

@@ -109,10 +109,12 @@ python3 -c "from harness.ledger import spent; print(spent())"
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
must fix, and he does not like Devstral. Devstral is not used further.
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
- Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0;
end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`.
The Devstral weights were deleted. No further base model tests.
- **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook,
run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`.
Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable.
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
## 7. Next steps