Files
abap-llm/train/README.md

4.8 KiB

Stage 1 training (train/)

Task: docs/stage1-training-task.md. State of the work: this file and train/STATE.md (later steps).

Environment

  • train/.venv: Python 3.11.17 (created with uv, Homebrew; the system Python 3.9 is not changed).
  • mlx-lm 0.32.0, mlx 0.32.3.

Base model

  • Base model: Qwen/Qwen3.8-27B (architecture qwen3_5, 27.8B, dense; Apache 2.0). It is the same base model as the Ollama model qwen3.8-27b-32k that the harness used before.
  • Weights: mlx-community/Qwen3.8-27B-4bit (Hugging Face), local path ~/models/Qwen3.8-27B-4bit. Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03.
  • Why not the Ollama weights (qwen3.8:27b-mlx): they are NVFP4 (modelopt) with one global scale per layer. mx.quantized_matmul has no global scale, so mlx_lm cannot load them without a re-quantization (a different model). Kral approved the Hugging Face download (2026-10-03).

Serving (train/serve.sh)

mlx_lm.server at http://127.0.0.1:8080/v1 (OpenAI-compatible). Base model without adapter; after training the same script with --adapter-path.

Settings, the same as the earlier Ollama runs (runs 103, 203):

Setting Ollama run mlx_lm.server
Thinking on, Ollama default level medium --chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}' (template default would be xhigh)
temperature 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) 0.2 (sent by the agent; server default --temp 0.2)
top_p / top_k / min_p 0.95 / 20 / 0 (Modelfile) --top-p 0.95 --top-k 20 --min-p 0
presence / repeat penalty 0 / 1 (neutral) not set (neutral)
Output limit none (context num_ctx 32768) --max-tokens 32768 (server default would be 512)
Context 32768 no fixed limit (memory)

Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about 100 visible tokens).

Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK) in the OpenAI tool_calls format; the arguments parse as JSON. The reasoning comes in a separate reasoning field. The chat template uses the qwen3_coder XML tool format; mlx_lm parses it. First request: 86 s (prompt of 9.5k tokens).

Eval subset

train/subset.json (copy: runs/stage1/subset.json): 11 tasks (Kral decision 2026-10-03): one task per category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (train/subset_v1_25.json) was cut because one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128, G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.

Settings of the stage 1 runs (the same for baseline and after training)

Setting Value
Model ~/models/Qwen3.8-27B-4bit (MLX affine 4 bit); after training the same with --adapter-path
Thinking off (enable_thinking: false), fixed for baseline and after training. Sent per request as chat_template_kwargs by train/baseline.py (the server default stays thinking on). Reason: with thinking on, all 3 baseline runs ended with empty responses at the thinking limit (runs/stage1/baseline_thinking_on.json)
temperature / top_p / top_k / min_p 0.2 / 0.95 / 20 / 0
max_tokens per turn 16384, sent in each request by train/baseline.py (MAX_TOKENS); the server limit stays 32768
Tool-call budget per task 60 calls, 15 activations (T01 too)
Loop guard loop_guard 3: the run ends when sap_push_source pushes the same source (object + md5) 3 times in a row; final report "Stopped: loop ...", end_reason "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it
Per run record end_reason (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), activation_failures, activation_error_messages (unique), also in baseline.json
Empty turn retried (2 times), then the run stops ("Stopped: empty model response")
Docker (A4H) VM memory 36 GB (MemoryMiB 36864), container --memory 32g --memory-swap 32g (2026-10-04)
Prompt cache of the server --prompt-cache-size 4 --prompt-cache-bytes 6000000000

The settings are also written into runs/stage1/baseline.json (settings). The old Ollama run (41.7), the T01 test with 32768 tokens and budget 40 (t01_test_budget40: 40.0) and the thinking-on runs are not comparable.

Baseline

  • Runner: python3 train/baseline.py --label baseline --run-base 20200 (one task at a time; results runs/stage1/baseline.json, run directories runs/stage1/baseline/). Started by train/baseline_chain.sh.