Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
10
CLAUDE.md
10
CLAUDE.md
@@ -102,6 +102,16 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
|||||||
Not yet run on SAP.
|
Not yet run on SAP.
|
||||||
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
|
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
|
||||||
|
|
||||||
|
## 6a. Base model decision (2026-10-04)
|
||||||
|
|
||||||
|
- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source
|
||||||
|
pushed again), empty responses at the thinking limit when thinking is on.
|
||||||
|
- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit
|
||||||
|
(`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`).
|
||||||
|
No thinking mode. All work is local; no Ollama cloud for this.
|
||||||
|
- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral.
|
||||||
|
- The qwen-specific lines in sections 2 and 6 are history.
|
||||||
|
|
||||||
## 7. Next steps
|
## 7. Next steps
|
||||||
|
|
||||||
1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`).
|
1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`).
|
||||||
|
|||||||
Reference in New Issue
Block a user