From 7351849a76f9961e7394e4dc2fe72b84030f4069 Mon Sep 17 00:00:00 2001 From: Kral Date: Sat, 3 Oct 2026 18:23:39 +0200 Subject: [PATCH] Stage 1 step 0: train/.venv (py3.11, mlx-lm 0.32), MLX 4-bit Qwen3.8-27B, serve script, baseline runner, README Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat --- train/README.md | 54 ++++++++++++++++++++++++++++++++++++++++++ train/baseline.py | 60 +++++++++++++++++++++++++++++++++++++++++++++++ train/serve.sh | 11 +++++++++ 3 files changed, 125 insertions(+) create mode 100644 train/README.md create mode 100644 train/baseline.py create mode 100755 train/serve.sh diff --git a/train/README.md b/train/README.md new file mode 100644 index 0000000..79f2c63 --- /dev/null +++ b/train/README.md @@ -0,0 +1,54 @@ +# Stage 1 training (train/) + +Task: `docs/stage1-training-task.md`. State of the work: this file and `train/STATE.md` (later steps). + +## Environment + +- `train/.venv`: Python 3.11.17 (created with `uv`, Homebrew; the system Python 3.9 is not changed). +- `mlx-lm` 0.32.0, `mlx` 0.32.3. + +## Base model + +- Base model: `Qwen/Qwen3.8-27B` (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). It is the same base + model as the Ollama model `qwen3.8-27b-32k` that the harness used before. +- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (Hugging Face), local path `~/models/Qwen3.8-27B-4bit`. + Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03. +- Why not the Ollama weights (`qwen3.8:27b-mlx`): they are NVFP4 (modelopt) with one global scale per + layer. `mx.quantized_matmul` has no global scale, so `mlx_lm` cannot load them without a re-quantization + (a different model). Kral approved the Hugging Face download (2026-10-03). + +## Serving (`train/serve.sh`) + +`mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after +training the same script with `--adapter-path`. + +Settings, the same as the earlier Ollama runs (runs 103, 203): + +| Setting | Ollama run | mlx_lm.server | +|---|---|---| +| Thinking | on, Ollama default level `medium` | `--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}'` (template default would be `xhigh`) | +| temperature | 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) | 0.2 (sent by the agent; server default `--temp 0.2`) | +| top_p / top_k / min_p | 0.95 / 20 / 0 (Modelfile) | `--top-p 0.95 --top-k 20 --min-p 0` | +| presence / repeat penalty | 0 / 1 (neutral) | not set (neutral) | +| Output limit | none (context `num_ctx` 32768) | `--max-tokens 32768` (server default would be 512) | +| Context | 32768 | no fixed limit (memory) | + +Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about +100 visible tokens). + +Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned +`sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK)` in the OpenAI `tool_calls` format; the +arguments parse as JSON. The reasoning comes in a separate `reasoning` field. The chat template uses the +qwen3_coder XML tool format; `mlx_lm` parses it. First request: 86 s (prompt of 9.5k tokens). + +## Eval subset + +`train/subset.json` (copy: `runs/stage1/subset.json`): 25 tasks, balanced over the categories +(A, B, C, E, F: 3 each; D, G, H, I, K: 2 each), with T01. Use the same list before and after training. + +## Baseline + +- Runner: `python3 train/baseline.py --label baseline` (one task at a time; results + `runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`). +- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 with the + MLX 4-bit weights is the new reference for stage 1. diff --git a/train/baseline.py b/train/baseline.py new file mode 100644 index 0000000..a628bf5 --- /dev/null +++ b/train/baseline.py @@ -0,0 +1,60 @@ +"""Stage 1 eval runs on the fixed subset (train/subset.json) with a local OpenAI-compatible server. + + train/.venv not needed: python3 train/baseline.py --label baseline [--only T01] [--run-base 20000] + +One task at a time. Results: runs/stage1/