Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
19
CLAUDE.md
19
CLAUDE.md
@@ -102,15 +102,18 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
|||||||
Not yet run on SAP.
|
Not yet run on SAP.
|
||||||
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
|
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
|
||||||
|
|
||||||
## 6a. Base model decision (2026-10-04)
|
## 6a. Base model decision (2026-10-04, revised)
|
||||||
|
|
||||||
- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source
|
- **Base model: Qwen 3.8 27B** (`Qwen/Qwen3.8-27B`, Apache 2.0; weights `mlx-community/Qwen3.8-27B-4bit`,
|
||||||
pushed again), empty responses at the thinking limit when thinking is on.
|
local `~/models/Qwen3.8-27B-4bit`). Kral decision, same day, after the Devstral baseline.
|
||||||
- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit
|
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
|
||||||
(`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`).
|
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
|
||||||
No thinking mode. All work is local; no Ollama cloud for this.
|
must fix, and he does not like Devstral. Devstral is not used further.
|
||||||
- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral.
|
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
|
||||||
- The qwen-specific lines in sections 2 and 6 are history.
|
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
|
||||||
|
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
|
||||||
|
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
|
||||||
|
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
|
||||||
|
|
||||||
## 7. Next steps
|
## 7. Next steps
|
||||||
|
|
||||||
|
|||||||
@@ -18,7 +18,7 @@ Durum: v20 · 2026-10-04
|
|||||||
| Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) |
|
| Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) |
|
||||||
| Lisans: Apache 2.0 | ✓ |
|
| Lisans: Apache 2.0 | ✓ |
|
||||||
| Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) |
|
| Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) |
|
||||||
| Baz model adayı: Devstral Small 2 (24B, Apache 2.0, MLX 4-bit). Qwen 3.8 elendi: aktivasyon hatasından sonra onarım yok, döngüler (aynı kaynak tekrar), thinking açıkken düşünme limitinde boş yanıt. Tüm iş yerel, Ollama cloud yok. Aşama 1 eğitim testi Devstral ile | ✓ (2026-10-04) |
|
| **Baz model: Qwen 3.8 27B** (Apache 2.0, MLX 4-bit). Aynı gün önce elendi (onarım yok, döngüler, thinking limitinde boş yanıt), Devstral Small 2 denendi (baseline 11 görevde ortalama 6,8). Kral kararıyla Qwen'e dönüldü: zayıflıklar eğitimin hedefi. Devstral kullanılmayacak. Aşama 1 eğitim testi Qwen ile | ✓ (2026-10-04) |
|
||||||
|
|
||||||
Bitti. ✓
|
Bitti. ✓
|
||||||
|
|
||||||
@@ -115,7 +115,7 @@ Kurallar:
|
|||||||
|
|
||||||
## Adım 2 — Baz model seçimi
|
## Adım 2 — Baz model seçimi
|
||||||
|
|
||||||
Durum (2026-10-04): aday Devstral Small 2. İlk baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Qwen elendi (`runs/archive/qwen38/`).
|
Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Devstral baseline'ı yalnız karşılaştırma. Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`).
|
||||||
|
|
||||||
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.
|
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.
|
||||||
|
|
||||||
|
|||||||
@@ -9,15 +9,18 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
|
|||||||
|
|
||||||
## Base model
|
## Base model
|
||||||
|
|
||||||
- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`,
|
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
|
||||||
24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode).
|
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
|
||||||
- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64,
|
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
|
||||||
15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for
|
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
|
||||||
the baseline, for training and for the run after training.
|
cannot load them without a re-quantization.
|
||||||
- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops
|
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
|
||||||
(same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results:
|
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
|
||||||
`runs/archive/qwen38/`. The Qwen sections below the baseline are history.
|
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
||||||
- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0).
|
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
||||||
|
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000).
|
||||||
|
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
|
||||||
|
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
|
||||||
|
|
||||||
## Serving (`train/serve.sh`)
|
## Serving (`train/serve.sh`)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user