From dbe6077a48fa66ac07fa9bf117286575e122c3cd Mon Sep 17 00:00:00 2001 From: Kral Date: Sun, 4 Oct 2026 13:11:53 +0200 Subject: [PATCH] Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further) Co-Authored-By: Claude Sonnet 5.5 Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat --- CLAUDE.md | 19 +++++++++++-------- docs/yol-haritasi.md | 4 ++-- train/README.md | 21 ++++++++++++--------- 3 files changed, 25 insertions(+), 19 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index cb1415e..6acdfb4 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -102,15 +102,18 @@ python3 -c "from harness.ledger import spent; print(spent())" Not yet run on SAP. - Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started. -## 6a. Base model decision (2026-10-04) +## 6a. Base model decision (2026-10-04, revised) -- Qwen 3.8 (27B) is dropped as base model. Reasons: no repair after activation errors, loops (same source - pushed again), empty responses at the thinking limit when thinking is on. -- New base candidate: Devstral Small 2 (24B, Apache 2.0), MLX 4-bit - (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, local `~/models/Devstral-Small-2-24B-4bit`). - No thinking mode. All work is local; no Ollama cloud for this. -- Qwen results: `runs/archive/qwen38/` (not deleted). Stage 1 baseline and training test now use Devstral. -- The qwen-specific lines in sections 2 and 6 are history. +- **Base model: Qwen 3.8 27B** (`Qwen/Qwen3.8-27B`, Apache 2.0; weights `mlx-community/Qwen3.8-27B-4bit`, + local `~/models/Qwen3.8-27B-4bit`). Kral decision, same day, after the Devstral baseline. +- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking + limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training + must fix, and he does not like Devstral. Devstral is not used further. +- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0. + `runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`. +- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the + MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference. +- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl). ## 7. Next steps diff --git a/docs/yol-haritasi.md b/docs/yol-haritasi.md index c112a40..0952330 100644 --- a/docs/yol-haritasi.md +++ b/docs/yol-haritasi.md @@ -18,7 +18,7 @@ Durum: v20 · 2026-10-04 | Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) | | Lisans: Apache 2.0 | ✓ | | Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) | -| Baz model adayı: Devstral Small 2 (24B, Apache 2.0, MLX 4-bit). Qwen 3.8 elendi: aktivasyon hatasından sonra onarım yok, döngüler (aynı kaynak tekrar), thinking açıkken düşünme limitinde boş yanıt. Tüm iş yerel, Ollama cloud yok. Aşama 1 eğitim testi Devstral ile | ✓ (2026-10-04) | +| **Baz model: Qwen 3.8 27B** (Apache 2.0, MLX 4-bit). Aynı gün önce elendi (onarım yok, döngüler, thinking limitinde boş yanıt), Devstral Small 2 denendi (baseline 11 görevde ortalama 6,8). Kral kararıyla Qwen'e dönüldü: zayıflıklar eğitimin hedefi. Devstral kullanılmayacak. Aşama 1 eğitim testi Qwen ile | ✓ (2026-10-04) | Bitti. ✓ @@ -115,7 +115,7 @@ Kurallar: ## Adım 2 — Baz model seçimi -Durum (2026-10-04): aday Devstral Small 2. İlk baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Qwen elendi (`runs/archive/qwen38/`). +Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Devstral baseline'ı yalnız karşılaştırma. Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`). Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik. diff --git a/train/README.md b/train/README.md index efb2bca..5347676 100644 --- a/train/README.md +++ b/train/README.md @@ -9,15 +9,18 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST ## Base model -- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`, - 24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode). -- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64, - 15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for - the baseline, for training and for the run after training. -- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops - (same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results: - `runs/archive/qwen38/`. The Qwen sections below the baseline are history. -- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0). +- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04 + (revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen). +- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path + `~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm` + cannot load them without a re-quantization. +- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`) + was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further. +- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), + empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1. +- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000). + `train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or + restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64). ## Serving (`train/serve.sh`)