Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
10
CLAUDE.md
10
CLAUDE.md
@@ -109,10 +109,12 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
||||
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
|
||||
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
|
||||
must fix, and he does not like Devstral. Devstral is not used further.
|
||||
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
|
||||
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
|
||||
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
|
||||
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
|
||||
- Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0;
|
||||
end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`.
|
||||
The Devstral weights were deleted. No further base model tests.
|
||||
- **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook,
|
||||
run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`.
|
||||
Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable.
|
||||
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
|
||||
|
||||
## 7. Next steps
|
||||
|
||||
Reference in New Issue
Block a user