Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 13:11:53 +02:00
parent 5532a5b226
commit dbe6077a48
3 changed files with 25 additions and 19 deletions

View File

@@ -9,15 +9,18 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
## Base model
- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`,
24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode).
- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64,
15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for
the baseline, for training and for the run after training.
- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops
(same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results:
`runs/archive/qwen38/`. The Qwen sections below the baseline are history.
- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0).
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
cannot load them without a re-quantization.
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000).
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
## Serving (`train/serve.sh`)