From 0af2d6458ebfcf0845916dc8aa6c946c2577b0be Mon Sep 17 00:00:00 2001 From: Kral Date: Sun, 4 Oct 2026 11:20:43 +0200 Subject: [PATCH] Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped Co-Authored-By: Claude Sonnet 5.5 Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat --- docs/yol-haritasi.md | 5 +++- train/README.md | 60 +++++++++++++++++++++++--------------------- 2 files changed, 36 insertions(+), 29 deletions(-) diff --git a/docs/yol-haritasi.md b/docs/yol-haritasi.md index 5d4fd09..c112a40 100644 --- a/docs/yol-haritasi.md +++ b/docs/yol-haritasi.md @@ -1,6 +1,6 @@ # ABAP Danışman Modeli — Yol Haritası -Durum: v19 · 2026-10-02 +Durum: v20 · 2026-10-04 Şu anki konum: **Adım 1.3** ## Adım 0 — Çerçeve ve kararlar @@ -18,6 +18,7 @@ Durum: v19 · 2026-10-02 | Eğitim verisinde Claude çıktısı yok; eval görevleri eğitim verisine girmez | ✓ (2026-10-02) | | Lisans: Apache 2.0 | ✓ | | Baz model: sadece Apache 2.0 veya MIT lisanslı adaylar | ✓ (seçim Adım 2'de) | +| Baz model adayı: Devstral Small 2 (24B, Apache 2.0, MLX 4-bit). Qwen 3.8 elendi: aktivasyon hatasından sonra onarım yok, döngüler (aynı kaynak tekrar), thinking açıkken düşünme limitinde boş yanıt. Tüm iş yerel, Ollama cloud yok. Aşama 1 eğitim testi Devstral ile | ✓ (2026-10-04) | Bitti. ✓ @@ -114,6 +115,8 @@ Kurallar: ## Adım 2 — Baz model seçimi +Durum (2026-10-04): aday Devstral Small 2. İlk baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Qwen elendi (`runs/archive/qwen38/`). + Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik. Bitti sayılır: tek baz model seçildi, gerekçesi yazıldı. diff --git a/train/README.md b/train/README.md index 7f842a2..efb2bca 100644 --- a/train/README.md +++ b/train/README.md @@ -9,37 +9,34 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST ## Base model -- Base model: `Qwen/Qwen3.8-27B` (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). It is the same base - model as the Ollama model `qwen3.8-27b-32k` that the harness used before. -- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (Hugging Face), local path `~/models/Qwen3.8-27B-4bit`. - Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03. -- Why not the Ollama weights (`qwen3.8:27b-mlx`): they are NVFP4 (modelopt) with one global scale per - layer. `mx.quantized_matmul` has no global scale, so `mlx_lm` cannot load them without a re-quantization - (a different model). Kral approved the Hugging Face download (2026-10-03). +- Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`, + 24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode). +- Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64, + 15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for + the baseline, for training and for the run after training. +- **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops + (same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results: + `runs/archive/qwen38/`. The Qwen sections below the baseline are history. +- Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0). ## Serving (`train/serve.sh`) `mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after training the same script with `--adapter-path`. -Settings, the same as the earlier Ollama runs (runs 103, 203): +| Setting | Value | +|---|---| +| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent | +| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) | +| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) | +| presence / repeat penalty | not set (neutral) | +| Output limit | server `--max-tokens 32768`; each request sends 16384 | +| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` | -| Setting | Ollama run | mlx_lm.server | -|---|---|---| -| Thinking | on, Ollama default level `medium` | `--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}'` (template default would be `xhigh`) | -| temperature | 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) | 0.2 (sent by the agent; server default `--temp 0.2`) | -| top_p / top_k / min_p | 0.95 / 20 / 0 (Modelfile) | `--top-p 0.95 --top-k 20 --min-p 0` | -| presence / repeat penalty | 0 / 1 (neutral) | not set (neutral) | -| Output limit | none (context `num_ctx` 32768) | `--max-tokens 32768` (server default would be 512) | -| Context | 32768 | no fixed limit (memory) | - -Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about -100 visible tokens). - -Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned -`sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK)` in the OpenAI `tool_calls` format; the -arguments parse as JSON. The reasoning comes in a separate `reasoning` field. The chat template uses the -qwen3_coder XML tool format; `mlx_lm` parses it. First request: 86 s (prompt of 9.5k tokens). +Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`); +`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors. +Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report +(G0128, G0174 ended with `report`). ## Eval subset @@ -52,8 +49,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra | Setting | Value | |---|---| -| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` | -| Thinking | **off (`enable_thinking: false`), fixed for baseline and after training.** Sent per request as `chat_template_kwargs` by `train/baseline.py` (the server default stays thinking on). Reason: with thinking on, all 3 baseline runs ended with empty responses at the thinking limit (`runs/stage1/baseline_thinking_on.json`) | +| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` | +| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) | | temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 | | max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 | | Tool-call budget per task | 60 calls, 15 activations (T01 too) | @@ -68,5 +65,12 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi ## Baseline -- Runner: `python3 train/baseline.py --label baseline --run-base 20200` (one task at a time; results - `runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`). Started by `train/baseline_chain.sh`. +- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by + `train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). + Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories + `runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`. +- Result: mean 6.8; 1 of 11 tasks above 0 (G0167: 75). End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 + (the model changes the source, but the changes do not remove the cause). +- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`. +- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`). +- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.