Files
abap-llm/train
Kral 301c221d8e Budget guard from panel 50.00: limit 128
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 21:31:15 +02:00
..
2026-10-04 17:02:59 +02:00

Stage 1 training (train/)

Task: docs/stage1-training-task.md. State of the work: this file and train/STATE.md (later steps).

Environment

  • train/.venv: Python 3.11.17 (created with uv, Homebrew; the system Python 3.9 is not changed).
  • mlx-lm 0.32.0, mlx 0.32.3.

Base model

  • Base model: Qwen/Qwen3.8-27B (architecture qwen3_5, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04 (Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
  • Weights: mlx-community/Qwen3.8-27B-4bit (MLX affine, 4 bit, group size 64, 16.1 GB), local path ~/models/Qwen3.8-27B-4bit. Why not the Ollama weights: they are NVFP4 with a global scale per layer; mlx_lm cannot load them without a re-quantization.
  • Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0, loops 6 vs 7. Results: runs/archive/devstral/. Weights deleted. No more base model tests.
  • Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
  • Stage 1 reference baseline: Qwen, mean 15.8, 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (runs/stage1/baseline_qwen.json, run base 22000; copy baseline.json). Comparison: docs/stage1-baseline.md. train/serve.sh serves Qwen (commit a2ba9e7).

Serving (train/serve.sh)

mlx_lm.server at http://127.0.0.1:8080/v1 (OpenAI-compatible). Base model without adapter; after training the same script with --adapter-path.

Setting Value
Thinking server default on (--chat-template-args); train/baseline.py sends enable_thinking=false per request
temperature 0.2 (sent by the harness llm agent; the model card suggests 0.15)
top_p / top_k / min_p 0.95 / 20 / 0 (server flags)
presence / repeat penalty not set (neutral)
Output limit server --max-tokens 32768; each request sends 16384
Prompt cache --prompt-cache-size 4 --prompt-cache-bytes 6000000000

Tool calls: mlx_lm returns OpenAI tool_calls with JSON arguments (tool-call test passed).

Eval subset

train/subset.json (copy: runs/stage1/subset.json): 11 tasks (Kral decision 2026-10-03): one task per category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (train/subset_v1_25.json) was cut because one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128, G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.

Settings of the stage 1 runs (the same for baseline and after training)

Setting Value
Model ~/models/Qwen3.8-27B-4bit (MLX affine 4 bit); after training the same with --adapter-path
Thinking off (enable_thinking=false per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (runs/archive/qwen38/baseline_thinking_on.json)
temperature / top_p / top_k / min_p 0.2 / 0.95 / 20 / 0
max_tokens per turn 16384, sent in each request by train/baseline.py (MAX_TOKENS); the server limit stays 32768
Tool-call budget per task 60 calls, 15 activations (T01 too)
Loop guard loop_guard 3: the run ends when sap_push_source pushes the same source (object + md5) 3 times in a row, or when any tool is called 3 times in a row with the same arguments and the same result (read loop, e.g. G0105 pulled the same include 18 times); final report "Stopped: loop ...", end_reason "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it
Per run record end_reason (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), activation_failures, activation_error_messages (unique), also in baseline.json
Empty turn retried (2 times), then the run stops ("Stopped: empty model response")
Docker (A4H) VM memory 36 GB (MemoryMiB 36864), container --memory 32g --memory-swap 32g (2026-10-04)
Prompt cache of the server --prompt-cache-size 4 --prompt-cache-bytes 6000000000

The settings are also written into runs/stage1/baseline.json (settings). The old Ollama run (41.7), the T01 test with 32768 tokens and budget 40 (t01_test_budget40: 40.0) and the thinking-on runs are not comparable.

Baseline

  • Official baseline: Qwen 3.8 27B, python3 train/baseline.py --label baseline_qwen --run-base 22000 (MacBook, thinking off). Results runs/stage1/baseline.json (= baseline_qwen.json), run directories runs/stage1/baseline_qwen/. Mean 15.8; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1. Repair rate 39/54 = 0.72.
  • Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84. Results in runs/archive/devstral/.
  • Per run record has end_reason, pushes_after_error, pushes_changed_after_error, repair_rate.
  • The stage 1 training test (step 3) uses Qwen (config_test.yaml).