Files
abap-llm/train

Stage 1 training (train/)

Task: docs/stage1-training-task.md. State of the work: this file and train/STATE.md (later steps).

Environment

  • train/.venv: Python 3.11.17 (created with uv, Homebrew; the system Python 3.9 is not changed).
  • mlx-lm 0.32.0, mlx 0.32.3.

Base model

  • Base model: Qwen/Qwen3.8-27B (architecture qwen3_5, 27.8B, dense; Apache 2.0). It is the same base model as the Ollama model qwen3.8-27b-32k that the harness used before.
  • Weights: mlx-community/Qwen3.8-27B-4bit (Hugging Face), local path ~/models/Qwen3.8-27B-4bit. Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03.
  • Why not the Ollama weights (qwen3.8:27b-mlx): they are NVFP4 (modelopt) with one global scale per layer. mx.quantized_matmul has no global scale, so mlx_lm cannot load them without a re-quantization (a different model). Kral approved the Hugging Face download (2026-10-03).

Serving (train/serve.sh)

mlx_lm.server at http://127.0.0.1:8080/v1 (OpenAI-compatible). Base model without adapter; after training the same script with --adapter-path.

Settings, the same as the earlier Ollama runs (runs 103, 203):

Setting Ollama run mlx_lm.server
Thinking on, Ollama default level medium --chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}' (template default would be xhigh)
temperature 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) 0.2 (sent by the agent; server default --temp 0.2)
top_p / top_k / min_p 0.95 / 20 / 0 (Modelfile) --top-p 0.95 --top-k 20 --min-p 0
presence / repeat penalty 0 / 1 (neutral) not set (neutral)
Output limit none (context num_ctx 32768) --max-tokens 32768 (server default would be 512)
Context 32768 no fixed limit (memory)

Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about 100 visible tokens).

Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK) in the OpenAI tool_calls format; the arguments parse as JSON. The reasoning comes in a separate reasoning field. The chat template uses the qwen3_coder XML tool format; mlx_lm parses it. First request: 86 s (prompt of 9.5k tokens).

Eval subset

train/subset.json (copy: runs/stage1/subset.json): 25 tasks, balanced over the categories (A, B, C, E, F: 3 each; D, G, H, I, K: 2 each), with T01. Use the same list before and after training.

Baseline

  • Runner: python3 train/baseline.py --label baseline (one task at a time; results runs/stage1/baseline.json, run directories runs/stage1/baseline/).
  • The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 with the MLX 4-bit weights is the new reference for stage 1.