Stage 1 step 0: train/.venv (py3.11, mlx-lm 0.32), MLX 4-bit Qwen3.8-27B, serve script, baseline runner, README

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 18:23:39 +02:00
parent cf8d6d5d07
commit 7351849a76
3 changed files with 125 additions and 0 deletions

54
train/README.md Normal file
View File

@@ -0,0 +1,54 @@
# Stage 1 training (train/)
Task: `docs/stage1-training-task.md`. State of the work: this file and `train/STATE.md` (later steps).
## Environment
- `train/.venv`: Python 3.11.17 (created with `uv`, Homebrew; the system Python 3.9 is not changed).
- `mlx-lm` 0.32.0, `mlx` 0.32.3.
## Base model
- Base model: `Qwen/Qwen3.8-27B` (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). It is the same base
model as the Ollama model `qwen3.8-27b-32k` that the harness used before.
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (Hugging Face), local path `~/models/Qwen3.8-27B-4bit`.
Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03.
- Why not the Ollama weights (`qwen3.8:27b-mlx`): they are NVFP4 (modelopt) with one global scale per
layer. `mx.quantized_matmul` has no global scale, so `mlx_lm` cannot load them without a re-quantization
(a different model). Kral approved the Hugging Face download (2026-10-03).
## Serving (`train/serve.sh`)
`mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after
training the same script with `--adapter-path`.
Settings, the same as the earlier Ollama runs (runs 103, 203):
| Setting | Ollama run | mlx_lm.server |
|---|---|---|
| Thinking | on, Ollama default level `medium` | `--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}'` (template default would be `xhigh`) |
| temperature | 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) | 0.2 (sent by the agent; server default `--temp 0.2`) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (Modelfile) | `--top-p 0.95 --top-k 20 --min-p 0` |
| presence / repeat penalty | 0 / 1 (neutral) | not set (neutral) |
| Output limit | none (context `num_ctx` 32768) | `--max-tokens 32768` (server default would be 512) |
| Context | 32768 | no fixed limit (memory) |
Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about
100 visible tokens).
Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned
`sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK)` in the OpenAI `tool_calls` format; the
arguments parse as JSON. The reasoning comes in a separate `reasoning` field. The chat template uses the
qwen3_coder XML tool format; `mlx_lm` parses it. First request: 86 s (prompt of 9.5k tokens).
## Eval subset
`train/subset.json` (copy: `runs/stage1/subset.json`): 25 tasks, balanced over the categories
(A, B, C, E, F: 3 each; D, G, H, I, K: 2 each), with T01. Use the same list before and after training.
## Baseline
- Runner: `python3 train/baseline.py --label baseline` (one task at a time; results
`runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`).
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 with the
MLX 4-bit weights is the new reference for stage 1.

60
train/baseline.py Normal file
View File

@@ -0,0 +1,60 @@
"""Stage 1 eval runs on the fixed subset (train/subset.json) with a local OpenAI-compatible server.
train/.venv not needed: python3 train/baseline.py --label baseline [--only T01] [--run-base 20000]
One task at a time. Results: runs/stage1/<label>.json (written after each task).
"""
import argparse
import json
import os
import sys
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
from harness.adt_client import load_env # noqa: E402
from harness.agents import LlmAgent # noqa: E402
from harness.runner import Runner # noqa: E402
MODEL = os.path.expanduser("~/models/Qwen3.8-27B-4bit")
BASE_URL = "http://127.0.0.1:8080/v1"
def main():
load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser()
ap.add_argument("--label", default="baseline")
ap.add_argument("--only", nargs="*")
ap.add_argument("--run-base", type=int, default=20000)
a = ap.parse_args()
subset = json.load(open(os.path.join(ROOT, "train", "subset.json")))["tasks"]
out_path = os.path.join(ROOT, "runs", "stage1", f"{a.label}.json")
os.makedirs(os.path.dirname(out_path), exist_ok=True)
res = json.load(open(out_path)) if os.path.exists(out_path) else {"model": MODEL, "tasks": {}}
runs_root = os.path.join(ROOT, "runs", "stage1", a.label)
for i, t in enumerate(subset):
tid = t["id"]
if (a.only and tid not in a.only) or tid in res["tasks"]:
continue
pool = os.path.join(ROOT, "tasks") if tid.startswith("T") else os.path.join(ROOT, "tasks_gen", "eval")
agent = LlmAgent(MODEL, BASE_URL)
try:
rep, run_dir = Runner(pool, runs_root).run(tid, agent, a.run_base + i)
except Exception as e: # noqa: BLE001 one broken run must not stop the series
res["tasks"][tid] = {"error": str(e)[:300]}
json.dump(res, open(out_path, "w"), indent=1)
print(json.dumps({"task": tid, "error": str(e)[:300]}), flush=True)
continue
h = rep.get("hidden_tests") or {}
res["tasks"][tid] = {"category": t["category"], "object_type": t["object_type"],
"score": (rep.get("score") or {}).get("total"), "parts": rep.get("score"),
"gates": rep.get("gates"), "hidden": f"{h.get('passed')}/{h.get('total')}",
"tool_calls": rep.get("tool_calls"), "seconds": rep.get("seconds"),
"agent_seconds": rep.get("agent_seconds"), "final": (rep.get("final_report") or "")[:300],
"run_dir": os.path.relpath(run_dir, ROOT)}
json.dump(res, open(out_path, "w"), indent=1)
print(json.dumps({"task": tid, "score": res["tasks"][tid]["score"], "hidden": res["tasks"][tid]["hidden"],
"tool_calls": rep.get("tool_calls"), "seconds": rep.get("seconds")}), flush=True)
if __name__ == "__main__":
main()

11
train/serve.sh Executable file
View File

@@ -0,0 +1,11 @@
#!/bin/sh
# Serve the base model (no adapter) with mlx_lm.server. Settings: train/README.md.
# Usage: train/serve.sh [--adapter-path PATH]
cd "$(dirname "$0")/.."
exec train/.venv/bin/mlx_lm.server \
--model "$HOME/models/Qwen3.8-27B-4bit" \
--host 127.0.0.1 --port 8080 \
--temp 0.2 --top-p 0.95 --top-k 20 --min-p 0 \
--max-tokens 32768 \
--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}' \
"$@"