Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
5.7 KiB
Stage 1 training (train/)
Task: docs/stage1-training-task.md. State of the work: this file and train/STATE.md (later steps).
Environment
train/.venv: Python 3.11.17 (created withuv, Homebrew; the system Python 3.9 is not changed).mlx-lm0.32.0,mlx0.32.3.
Base model
- Base model:
Qwen/Qwen3.8-27B(architectureqwen3_5, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04 (revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen). - Weights:
mlx-community/Qwen3.8-27B-4bit(MLX affine, 4 bit, group size 64, 16.1 GB), local path~/models/Qwen3.8-27B-4bit. Why not the Ollama weights: they are NVFP4 with a global scale per layer;mlx_lmcannot load them without a re-quantization. - Devstral Small 2 (
mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit,~/models/Devstral-Small-2-24B-4bit) was tested only for the baseline:runs/stage1/baseline_devstral.md(mean 6.8, 1 of 11 tasks above 0). Not used further. - Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (
baseline_qwen.json, run base 22000).train/serve.shserves Devstral at the moment; for Qwen usetrain/serve_qwen.sh(in the MacBook package) or restore the Qwen line (--model ~/models/Qwen3.8-27B-4bit,--chat-template-argsas in git history before0af2d64).
Serving (train/serve.sh)
mlx_lm.server at http://127.0.0.1:8080/v1 (OpenAI-compatible). Base model without adapter; after
training the same script with --adapter-path.
| Setting | Value |
|---|---|
| Thinking | none (Devstral has no thinking mode); no chat_template_kwargs are sent |
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
| presence / repeat penalty | not set (neutral) |
| Output limit | server --max-tokens 32768; each request sends 16384 |
| Prompt cache | --prompt-cache-size 4 --prompt-cache-bytes 6000000000 |
Tool calls (2026-10-04): the chat template uses the Mistral format ([AVAILABLE_TOOLS], [TOOL_CALLS]name[ARGS]{json});
mlx_lm returns OpenAI tool_calls with JSON arguments. The smoke test T01 and the baseline had no parse errors.
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
(G0128, G0174 ended with report).
Eval subset
train/subset.json (copy: runs/stage1/subset.json): 11 tasks (Kral decision 2026-10-03): one task per
category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (train/subset_v1_25.json) was cut because
one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128,
G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.
Settings of the stage 1 runs (the same for baseline and after training)
| Setting | Value |
|---|---|
| Model | ~/models/Devstral-Small-2-24B-4bit (MLX affine 4 bit); after training the same with --adapter-path |
| Thinking | none: Devstral has no thinking mode; no enable_thinking is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, runs/archive/qwen38/baseline_thinking_on.json) |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | 16384, sent in each request by train/baseline.py (MAX_TOKENS); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
| Loop guard | loop_guard 3: the run ends when sap_push_source pushes the same source (object + md5) 3 times in a row, or when any tool is called 3 times in a row with the same arguments and the same result (read loop, e.g. G0105 pulled the same include 18 times); final report "Stopped: loop ...", end_reason "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it |
| Per run record | end_reason (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), activation_failures, activation_error_messages (unique), also in baseline.json |
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
| Docker (A4H) | VM memory 36 GB (MemoryMiB 36864), container --memory 32g --memory-swap 32g (2026-10-04) |
| Prompt cache of the server | --prompt-cache-size 4 --prompt-cache-bytes 6000000000 |
The settings are also written into runs/stage1/baseline.json (settings). The old Ollama run (41.7), the
T01 test with 32768 tokens and budget 40 (t01_test_budget40: 40.0) and the thinking-on runs are not comparable.
Baseline
- Devstral baseline (2026-10-04):
python3 train/baseline.py --label baseline_devstral --run-base 21000, started bytrain/baseline_chain.sh(stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). Resultsruns/stage1/baseline_devstral.json(copy:runs/stage1/baseline.json), run directoriesruns/stage1/baseline_devstral/, reportruns/stage1/baseline_devstral.md. - Result: mean 6.8; 1 of 11 tasks above 0 (G0167: 75). End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 (the model changes the source, but the changes do not remove the cause).
- Per run record now also has
pushes_after_error,pushes_changed_after_error,repair_rate. - The Qwen baseline was never completed (archive:
runs/archive/qwen38/). - The stage 1 training test (step 3) now uses Devstral.
mlx_lm.loramust be checked formistral3before the test.