# Stage 1 training (train/) Task: `docs/stage1-training-task.md`. State of the work: this file and `train/STATE.md` (later steps). ## Environment - `train/.venv`: Python 3.11.17 (created with `uv`, Homebrew; the system Python 3.9 is not changed). - `mlx-lm` 0.32.0, `mlx` 0.32.3. ## Base model - Base model (candidate, 2026-10-04): **Devstral Small 2** (`mistralai/Devstral-Small-2-24B-Instruct-2512`, 24B dense, `Mistral3ForConditionalGeneration`, Apache 2.0; no thinking mode). - Weights: **`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`** (MLX affine, 4 bit, group size 64, 15.1 GB, text and vision tower), local path `~/models/Devstral-Small-2-24B-4bit`. The same build is used for the baseline, for training and for the run after training. - **Qwen 3.8 (27B) is dropped** (Kral decision 2026-10-04). Reasons: no repair after activation errors, loops (same source pushed again), empty responses at the thinking limit when thinking is on. Qwen results: `runs/archive/qwen38/`. The Qwen sections below the baseline are history. - Result of the first Devstral baseline: `runs/stage1/baseline_devstral.md` (1 of 11 tasks above 0). ## Serving (`train/serve.sh`) `mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after training the same script with `--adapter-path`. | Setting | Value | |---|---| | Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent | | temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) | | top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) | | presence / repeat penalty | not set (neutral) | | Output limit | server `--max-tokens 32768`; each request sends 16384 | | Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` | Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`); `mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors. Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report (G0128, G0174 ended with `report`). ## Eval subset `train/subset.json` (copy: `runs/stage1/subset.json`): **11 tasks** (Kral decision 2026-10-03): one task per category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (`train/subset_v1_25.json`) was cut because one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128, G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training. ## Settings of the stage 1 runs (the same for baseline and after training) | Setting | Value | |---|---| | Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` | | Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) | | temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 | | max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 | | Tool-call budget per task | 60 calls, 15 activations (T01 too) | | Loop guard | `loop_guard` 3: the run ends when `sap_push_source` pushes the same source (object + md5) 3 times in a row, or when any tool is called 3 times in a row with the same arguments and the same result (read loop, e.g. G0105 pulled the same include 18 times); final report "Stopped: loop ...", `end_reason` "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it | | Per run record | `end_reason` (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), `activation_failures`, `activation_error_messages` (unique), also in `baseline.json` | | Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") | | Docker (A4H) | VM memory 36 GB (`MemoryMiB` 36864), container `--memory 32g --memory-swap 32g` (2026-10-04) | | Prompt cache of the server | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` | The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7), the T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thinking-on runs are not comparable. ## Baseline - Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by `train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories `runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`. - Result: mean 6.8; 1 of 11 tasks above 0 (G0167: 75). End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 (the model changes the source, but the changes do not remove the cause). - Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`. - The Qwen baseline was never completed (archive: `runs/archive/qwen38/`). - The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.