Files
abap-llm/docs/local-qwen-run.md

5.5 KiB

Series A: local Qwen 3.8 27B on the new kinds (2026-10-05 22:24 to 2026-10-06 04:00)

Setup: docs/remote-model.md. Model served on the MacBook (/Users/I301710/models/Qwen3.8-27B-4bit, mlx-lm 0.32.0), harness, A4H and MCP on the Mac mini. Settings as the official baseline: thinking off, max_tokens 16384, loop guard 3, tool budget 60, temperature 0.2, same system prompt. Acceptance: score at least 80, end reason report. One task at a time. 20 tasks in 4.7 hours of run time (the 12 hour window was not used up); no server pause, no infrastructure outage, no harness error.

Result per kind

kind accepted / runs mean score failure types
INTF 1/3 33.3 loop 2, not_active 2, contract 2, hidden_tests_not_run 2
TABL 0/3 0.0 contract 3, hidden_tests_not_run 3, loop 2, not_active 2, tool_budget 1
STRU 0/3 25.0 loop 2, not_active 2, contract 2, hidden_tests_not_run 2, tool_budget 1
MSAG 2/3 66.7 tool_budget 1, not_active 1, contract 1, hidden_tests_not_run 1
EXC 0/3 0.0 loop 3, not_active 3, contract 3, hidden_tests_not_run 3
DDLS 0/5 31.0 loop 3, not_active 3, contract 3, hidden_tests_not_run 3, tool_budget 1, low_score 1

Total: 3 of 20 accepted (15 %). Failure types over all runs: {'loop': 12, 'not_active': 13, 'contract': 14, 'hidden_tests_not_run': 14, 'tool_budget': 4, 'low_score': 1}.

Reading

  • Loops are the failure. 12 of the 17 failed runs ended with end_reason loop (9 times the same source pushed three times in a row, 3 times the same read call with the same result). 10 of the 12 are short (1 to 6 minutes, 7 to 24 calls), two are long (59 and 71 minutes, 41 and 29 calls): Qwen pushes a source, gets an activation or save error and pushes the same source again instead of repairing it. This is the weakness the baseline showed (loop 7 of 11) and the main target of the SFT. Typical errors it did not repair: The component C_ZONE has been declared multiple times, The statement 50 is unexpected, Syntax error in <table>: DDL source could not be saved, FRIENDS ... expected after GLOBAL, Test method can be defined only in test classes.
  • Tool budget (60 calls), 4 runs, two different patterns the loop guard does not catch (the arguments differ each time): three runs did nothing but sap_search_object (STRU G1947: 60 searches, MSAG G1940: 60, DDLS G1012: 56 searches, no write at all), and TABL G1901 pushed 49 different sources in a row (51 writes, never active). The loop guard only sees identical calls; these are search and rewrite loops with changing arguments. A teacher trajectory that stops searching and writes, and one that repairs after the first activation error, are what the SFT must show.
  • Only INTF G1917 (10 calls) and MSAG G1902 and G1945 (13 and 19 calls) were solved cleanly: the small, regular tasks.
  • Two runs reached 75 to 80 points but not an accepted state (STRU G1903 loop at the end, DDLS G1011 loop, DDLS G1031 low score after 56 calls and 64 minutes).
  • Proxy syntax hint (local abaplint) was added in 12 calls (STRU G1903: 9, EXC G1124: 3). It did not help here: both runs still looped (the hint names a line and a parser message, Qwen pushed the same source again).
  • Speed: mean 14 minutes per run, median 5; the slow ones (STRU G1903 71 min, DDLS G1031 64 min, TABL G1910 59 min) are long generations, not pauses.
  • Compare with DeepSeek on the same kinds: see the pipeline summaries in train/STATE.md (teacher acceptance about 76 %).

Per task

task kind accepted score end calls minutes failure types
G1900 INTF no 0 loop 7 6 loop, not_active, contract, hidden_tests_not_run
G1913 INTF no 0 loop 8 5 loop, not_active, contract, hidden_tests_not_run
G1917 INTF accepted 100.0 report 10 7 -
G1901 TABL no 0 tool_budget 60 15 tool_budget, contract, hidden_tests_not_run
G1910 TABL no 0 loop 41 59 loop, not_active, contract, hidden_tests_not_run
G1911 TABL no 0 loop 15 2 loop, not_active, contract, hidden_tests_not_run
G1903 STRU no 75.0 loop 29 71 loop
G1946 STRU no 0 loop 24 4 loop, not_active, contract, hidden_tests_not_run
G1947 STRU no 0 tool_budget 60 5 tool_budget, not_active, contract, hidden_tests_not_run
G1902 MSAG accepted 100.0 report 19 9 -
G1940 MSAG no 0 tool_budget 60 6 tool_budget, not_active, contract, hidden_tests_not_run
G1945 MSAG accepted 100.0 report 13 3 -
G1096 EXC no 0 loop 14 3 loop, not_active, contract, hidden_tests_not_run
G1104 EXC no 0 loop 15 5 loop, not_active, contract, hidden_tests_not_run
G1124 EXC no 0 loop 8 1 loop, not_active, contract, hidden_tests_not_run
G1000 DDLS no 0 loop 8 2 loop, not_active, contract, hidden_tests_not_run
G1011 DDLS no 80.0 loop 18 4 loop
G1012 DDLS no 0 tool_budget 60 5 tool_budget, not_active, contract, hidden_tests_not_run
G1020 DDLS no 0 loop 11 3 loop, not_active, contract, hidden_tests_not_run
G1031 DDLS no 75.0 report 56 64 low_score

Files

runs/local_qwen/ (not in git): results.jsonl, summary.json, state.json, preflight.json, runs/ (records), and accepted.jsonl with the 3 accepted trajectories (metadata.use_for_sft = false, series local_qwen_A): not for the first SFT.