Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 11:20:43 +02:00
parent 8d1db9c67c
commit 0af2d6458e
2 changed files with 36 additions and 29 deletions

View File

@@ -1,6 +1,6 @@
# ABAP Danışman Modeli — Yol Haritası
Durum: v19 · 2026-10-02
Durum: v20 · 2026-10-04
Şu anki konum: **Adım 1.3**
## Adım 0 — Çerçeve ve kararlar
@@ -18,6 +18,7 @@ Durum: v19 · 2026-10-02
| Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) |
| Lisans: Apache 2.0 | ✓ |
| Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) |
| Baz model adayı: Devstral Small 2 (24B, Apache 2.0, MLX 4-bit). Qwen 3.8 elendi: aktivasyon hatasından sonra onarım yok, döngüler (aynı kaynak tekrar), thinking açıkken düşünme limitinde boş yanıt. Tüm iş yerel, Ollama cloud yok. Aşama 1 eğitim testi Devstral ile | ✓ (2026-10-04) |
Bitti. ✓
@@ -114,6 +115,8 @@ Kurallar:
## Adım 2 — Baz model seçimi
Durum (2026-10-04): aday Devstral Small 2. İlk baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Qwen elendi (`runs/archive/qwen38/`).
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.
Bitti sayılır: tek baz model seçildi, gerekçesi yazıldı.

View File

@@ -9,37 +9,34 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
## Base model
- Base model: `Qwen/Qwen3.8-27B` (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). It is the same base
model as the Ollama model `qwen3.8-27b-32k` that the harness used before.
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (Hugging Face), local path `~/models/Qwen3.8-27B-4bit`.
Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03.
- Why not the Ollama weights (`qwen3.8:27b-mlx`): they are NVFP4 (modelopt) with one global scale per
layer. `mx.quantized_matmul` has no global scale, so `mlx_lm` cannot load them without a re-quantization
(a different model). Kral approved the Hugging Face download (2026-10-03).
- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`,
24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode).
- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64,
15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for
the baseline, for training and for the run after training.
- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops
(same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results:
`runs/archive/qwen38/`. The Qwen sections below the baseline are history.
- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0).
## Serving (`train/serve.sh`)
`mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after
training the same script with `--adapter-path`.
Settings, the same as the earlier Ollama runs (runs 103, 203):
| Setting | Value |
|---|---|
| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent |
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
| presence / repeat penalty | not set (neutral) |
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
| Setting | Ollama run | mlx_lm.server |
|---|---|---|
| Thinking | on, Ollama default level `medium` | `--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}'` (template default would be `xhigh`) |
| temperature | 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) | 0.2 (sent by the agent; server default `--temp 0.2`) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (Modelfile) | `--top-p 0.95 --top-k 20 --min-p 0` |
| presence / repeat penalty | 0 / 1 (neutral) | not set (neutral) |
| Output limit | none (context `num_ctx` 32768) | `--max-tokens 32768` (server default would be 512) |
| Context | 32768 | no fixed limit (memory) |
Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about
100 visible tokens).
Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned
`sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK)` in the OpenAI `tool_calls` format; the
arguments parse as JSON. The reasoning comes in a separate `reasoning` field. The chat template uses the
qwen3_coder XML tool format; `mlx_lm` parses it. First request: 86 s (prompt of 9.5k tokens).
Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`);
`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors.
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
(G0128, G0174 ended with `report`).
## Eval subset
@@ -52,8 +49,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
| Setting | Value |
|---|---|
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | **off (`enable_thinking: false`), fixed for baseline and after training.** Sent per request as `chat_template_kwargs` by `train/baseline.py` (the server default stays thinking on). Reason: with thinking on, all 3 baseline runs ended with empty responses at the thinking limit (`runs/stage1/baseline_thinking_on.json`) |
| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
@@ -68,5 +65,12 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
## Baseline
- Runner: `python3 train/baseline.py --label baseline --run-base 20200` (one task at a time; results
`runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`). Started by `train/baseline_chain.sh`.
- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
- Result: mean 6.8; 1 of 11 tasks above 0 (G0167: 75). End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
(the model changes the source, but the changes do not remove the cause).
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).
- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.