Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 13:11:53 +02:00
parent 5532a5b226
commit dbe6077a48
3 changed files with 25 additions and 19 deletions

View File

@@ -102,15 +102,18 @@ python3 -c "from harness.ledger import spent; print(spent())"
Not yet run on SAP. Not yet run on SAP.
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started. - Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
## 6a. Base model decision (2026-10-04) ## 6a. Base model decision (2026-10-04, revised)
- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source - **Base model: Qwen 3.8 27B** (`Qwen/Qwen3.8-27B`, Apache 2.0; weights `mlx-community/Qwen3.8-27B-4bit`,
pushed again), empty responses at the thinking limit when thinking is on. local `~/models/Qwen3.8-27B-4bit`). Kral decision, same day, after the Devstral baseline.
- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit - Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
(`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`). limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
No thinking mode. All work is local; no Ollama cloud for this. must fix, and he does not like Devstral. Devstral is not used further.
- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral. - Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
- The qwen-specific lines in sections 2 and 6 are history. `runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
## 7. Next steps ## 7. Next steps

View File

@@ -18,7 +18,7 @@ Durum: v20 · 2026-10-04
| Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) | | Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) |
| Lisans: Apache 2.0 | ✓ | | Lisans: Apache 2.0 | ✓ |
| Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) | | Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) |
| Baz model adayı: Devstral Small 2 (24B, Apache 2.0, MLX 4-bit). Qwen 3.8 elendi: aktivasyon hatasından sonra onarım yok, döngüler (aynı kaynak tekrar), thinking açıkken düşünme limitinde boş yanıt. Tüm iş yerel, Ollama cloud yok. Aşama 1 eğitim testi Devstral ile | ✓ (2026-10-04) | | **Baz model: Qwen 3.8 27B** (Apache 2.0, MLX 4-bit). Aynı gün önce elendi (onarım yok, döngüler, thinking limitinde boş yanıt), Devstral Small 2 denendi (baseline 11 görevde ortalama 6,8). Kral kararıyla Qwen'e dönüldü: zayıflıklar eğitimin hedefi. Devstral kullanılmayacak. Aşama 1 eğitim testi Qwen ile | ✓ (2026-10-04) |
Bitti. ✓ Bitti. ✓
@@ -115,7 +115,7 @@ Kurallar:
## Adım 2 — Baz model seçimi ## Adım 2 — Baz model seçimi
Durum (2026-10-04): aday Devstral Small 2. İlk baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Qwen elendi (`runs/archive/qwen38/`). Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Devstral baseline'ı yalnız karşılaştırma. Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`).
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik. Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.

View File

@@ -9,15 +9,18 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
## Base model ## Base model
- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`, - Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode). (revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64, - Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for `~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
the baseline, for training and for the run after training. cannot load them without a re-quantization.
- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops - Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
(same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results: was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
`runs/archive/qwen38/`. The Qwen sections below the baseline are history. - Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0). empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000).
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
## Serving (`train/serve.sh`) ## Serving (`train/serve.sh`)