Files
abap-llm/train/README.md

5.9 KiB

Stage 1 training (train/)

Task: docs/stage1-training-task.md. State of the work: this file and train/STATE.md (later steps).

Environment

  • train/.venv: Python 3.11.17 (created with uv, Homebrew; the system Python 3.9 is not changed).
  • mlx-lm 0.32.0, mlx 0.32.3.

Base model

  • Base model: Qwen/Qwen3.8-27B (architecture qwen3_5, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04 (revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
  • Weights: mlx-community/Qwen3.8-27B-4bit (MLX affine, 4 bit, group size 64, 16.1 GB), local path ~/models/Qwen3.8-27B-4bit. Why not the Ollama weights: they are NVFP4 with a global scale per layer; mlx_lm cannot load them without a re-quantization.
  • Devstral Small 2 (mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit, ~/models/Devstral-Small-2-24B-4bit) was tested only for the baseline: runs/stage1/baseline_devstral.md (mean 6.8, 1 of 11 tasks above 0). Not used further.
  • Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
  • Stage 1 reference baseline: Qwen, mean 15.8, 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (runs/stage1/baseline_qwen.json, run base 22000; copy baseline.json). Comparison: docs/stage1-baseline.md. train/serve.sh serves Devstral at the moment; for Qwen use train/serve_qwen.sh (in the MacBook package) or restore the Qwen line (--model ~/models/Qwen3.8-27B-4bit, --chat-template-args as in git history before 0af2d64).

Serving (train/serve.sh)

mlx_lm.server at http://127.0.0.1:8080/v1 (OpenAI-compatible). Base model without adapter; after training the same script with --adapter-path.

Setting Value
Thinking none (Devstral has no thinking mode); no chat_template_kwargs are sent
temperature 0.2 (sent by the harness llm agent; the model card suggests 0.15)
top_p / top_k / min_p 0.95 / 20 / 0 (server flags)
presence / repeat penalty not set (neutral)
Output limit server --max-tokens 32768; each request sends 16384
Prompt cache --prompt-cache-size 4 --prompt-cache-bytes 6000000000

Tool calls (2026-10-04): the chat template uses the Mistral format ([AVAILABLE_TOOLS], [TOOL_CALLS]name[ARGS]{json}); mlx_lm returns OpenAI tool_calls with JSON arguments. The smoke test T01 and the baseline had no parse errors. Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report (G0128, G0174 ended with report).

Eval subset

train/subset.json (copy: runs/stage1/subset.json): 11 tasks (Kral decision 2026-10-03): one task per category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (train/subset_v1_25.json) was cut because one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128, G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.

Settings of the stage 1 runs (the same for baseline and after training)

Setting Value
Model ~/models/Devstral-Small-2-24B-4bit (MLX affine 4 bit); after training the same with --adapter-path
Thinking none: Devstral has no thinking mode; no enable_thinking is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, runs/archive/qwen38/baseline_thinking_on.json)
temperature / top_p / top_k / min_p 0.2 / 0.95 / 20 / 0
max_tokens per turn 16384, sent in each request by train/baseline.py (MAX_TOKENS); the server limit stays 32768
Tool-call budget per task 60 calls, 15 activations (T01 too)
Loop guard loop_guard 3: the run ends when sap_push_source pushes the same source (object + md5) 3 times in a row, or when any tool is called 3 times in a row with the same arguments and the same result (read loop, e.g. G0105 pulled the same include 18 times); final report "Stopped: loop ...", end_reason "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it
Per run record end_reason (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), activation_failures, activation_error_messages (unique), also in baseline.json
Empty turn retried (2 times), then the run stops ("Stopped: empty model response")
Docker (A4H) VM memory 36 GB (MemoryMiB 36864), container --memory 32g --memory-swap 32g (2026-10-04)
Prompt cache of the server --prompt-cache-size 4 --prompt-cache-bytes 6000000000

The settings are also written into runs/stage1/baseline.json (settings). The old Ollama run (41.7), the T01 test with 32768 tokens and budget 40 (t01_test_budget40: 40.0) and the thinking-on runs are not comparable.

Baseline

  • Devstral baseline (2026-10-04): python3 train/baseline.py --label baseline_devstral --run-base 21000, started by train/baseline_chain.sh (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). Results runs/stage1/baseline_devstral.json (copy: runs/stage1/baseline.json), run directories runs/stage1/baseline_devstral/, report runs/stage1/baseline_devstral.md.
  • Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 (the model changes the source, but the changes do not remove the cause).
  • Per run record now also has pushes_after_error, pushes_changed_after_error, repair_rate.
  • The Qwen baseline was never completed (archive: runs/archive/qwen38/).
  • The stage 1 training test (step 3) now uses Devstral. mlx_lm.lora must be checked for mistral3 before the test.