Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -9,15 +9,18 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
|
||||
|
||||
## Base model
|
||||
|
||||
- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`,
|
||||
24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode).
|
||||
- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64,
|
||||
15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for
|
||||
the baseline, for training and for the run after training.
|
||||
- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops
|
||||
(same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results:
|
||||
`runs/archive/qwen38/`. The Qwen sections below the baseline are history.
|
||||
- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0).
|
||||
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
|
||||
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
|
||||
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
|
||||
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
|
||||
cannot load them without a re-quantization.
|
||||
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
|
||||
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
|
||||
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
||||
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
||||
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000).
|
||||
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
|
||||
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
|
||||
|
||||
## Serving (`train/serve.sh`)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user