Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 08:44:04 +02:00
parent ba0a7e6b3b
commit e84ea43a3d

View File

@@ -102,6 +102,16 @@ python3 -c "from harness.ledger import spent; print(spent())"
Not yet run on SAP. Not yet run on SAP.
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started. - Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
## 6a. Base model decision (2026-10-04)
- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source
pushed again), empty responses at the thinking limit when thinking is on.
- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit
(`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`).
No thinking mode. All work is local; no Ollama cloud for this.
- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral.
- The qwen-specific lines in sections 2 and 6 are history.
## 7. Next steps ## 7. Next steps
1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`). 1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`).