From e84ea43a3d3bfb41afed8d4d2f0b58abc1b14401 Mon Sep 17 00:00:00 2001 From: Kral Date: Sun, 4 Oct 2026 08:44:04 +0200 Subject: [PATCH] Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B) Co-Authored-By: Claude Sonnet 5.5 Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat --- CLAUDE.md | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/CLAUDE.md b/CLAUDE.md index f97c7ab..cb1415e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -102,6 +102,16 @@ python3 -c "from harness.ledger import spent; print(spent())" Not yet run on SAP. - Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started. +## 6a. Base model decision (2026-10-04) + +- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source + pushed again), empty responses at the thinking limit when thinking is on. +- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit + (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`). + No thinking mode. All work is local; no Ollama cloud for this. +- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral. +- The qwen-specific lines in sections 2 and 6 are history. + ## 7. Next steps 1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`).