Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
19
CLAUDE.md
19
CLAUDE.md
@@ -102,15 +102,18 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
||||
Not yet run on SAP.
|
||||
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
|
||||
|
||||
## 6a. Base model decision (2026-10-04)
|
||||
## 6a. Base model decision (2026-10-04, revised)
|
||||
|
||||
- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source
|
||||
pushed again), empty responses at the thinking limit when thinking is on.
|
||||
- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit
|
||||
(`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`).
|
||||
No thinking mode. All work is local; no Ollama cloud for this.
|
||||
- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral.
|
||||
- The qwen-specific lines in sections 2 and 6 are history.
|
||||
- **Base model: Qwen 3.8 27B** (`Qwen/Qwen3.8-27B`, Apache 2.0; weights `mlx-community/Qwen3.8-27B-4bit`,
|
||||
local `~/models/Qwen3.8-27B-4bit`). Kral decision, same day, after the Devstral baseline.
|
||||
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
|
||||
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
|
||||
must fix, and he does not like Devstral. Devstral is not used further.
|
||||
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
|
||||
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
|
||||
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
|
||||
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
|
||||
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
|
||||
|
||||
## 7. Next steps
|
||||
|
||||
|
||||
Reference in New Issue
Block a user