Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
Stage 1 training (train/)
Task: docs/stage1-training-task.md. State of the work: this file and train/STATE.md (later steps).
Environment
train/.venv: Python 3.11.17 (created withuv, Homebrew; the system Python 3.9 is not changed).mlx-lm0.32.0,mlx0.32.3.
Base model
- Base model:
Qwen/Qwen3.8-27B(architectureqwen3_5, 27.8B, dense; Apache 2.0). It is the same base model as the Ollama modelqwen3.8-27b-32kthat the harness used before. - Weights:
mlx-community/Qwen3.8-27B-4bit(Hugging Face), local path~/models/Qwen3.8-27B-4bit. Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03. - Why not the Ollama weights (
qwen3.8:27b-mlx): they are NVFP4 (modelopt) with one global scale per layer.mx.quantized_matmulhas no global scale, somlx_lmcannot load them without a re-quantization (a different model). Kral approved the Hugging Face download (2026-10-03).
Serving (train/serve.sh)
mlx_lm.server at http://127.0.0.1:8080/v1 (OpenAI-compatible). Base model without adapter; after
training the same script with --adapter-path.
Settings, the same as the earlier Ollama runs (runs 103, 203):
| Setting | Ollama run | mlx_lm.server |
|---|---|---|
| Thinking | on, Ollama default level medium |
--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}' (template default would be xhigh) |
| temperature | 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) | 0.2 (sent by the agent; server default --temp 0.2) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (Modelfile) | --top-p 0.95 --top-k 20 --min-p 0 |
| presence / repeat penalty | 0 / 1 (neutral) | not set (neutral) |
| Output limit | none (context num_ctx 32768) |
--max-tokens 32768 (server default would be 512) |
| Context | 32768 | no fixed limit (memory) |
Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about 100 visible tokens).
Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned
sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK) in the OpenAI tool_calls format; the
arguments parse as JSON. The reasoning comes in a separate reasoning field. The chat template uses the
qwen3_coder XML tool format; mlx_lm parses it. First request: 86 s (prompt of 9.5k tokens).
Eval subset
train/subset.json (copy: runs/stage1/subset.json): 11 tasks (Kral decision 2026-10-03): one task per
category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (train/subset_v1_25.json) was cut because
one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128,
G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.
Settings of the stage 1 runs (the same for baseline and after training)
| Setting | Value |
|---|---|
| Model | ~/models/Qwen3.8-27B-4bit (MLX affine 4 bit); after training the same with --adapter-path |
| Thinking | on, reasoning_effort medium |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | 16384, sent in each request by train/baseline.py (MAX_TOKENS); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
| Prompt cache of the server | --prompt-cache-size 4 --prompt-cache-bytes 6000000000 |
The settings are also written into runs/stage1/baseline.json (settings). The old Ollama run (41.7) and the
T01 test with 32768 tokens and budget 40 (t01_test_budget40: 40.0) are not comparable.
Baseline
- Runner:
python3 train/baseline.py --label baseline --run-base 20200(one task at a time; resultsruns/stage1/baseline.json, run directoriesruns/stage1/baseline/). Started bytrain/baseline_chain.sh.