Files
abap-llm/docs/stage1-baseline.md

3.9 KiB
Raw Blame History

Stage 1 baseline: Qwen 3.8 27B vs Devstral Small 2

Date: 2026-10-04. Subset: train/subset.json (11 tasks). Same settings for both: no thinking (Qwen: enable_thinking=false per request; Devstral has no thinking mode), temperature 0.2, max_tokens 16384, tool-call budget 60, loop guard 3. Both models: MLX affine 4 bit (Qwen3.8-27B-4bit, Devstral-Small-2-24B-4bit). Qwen ran on the MacBook (M4 Pro 48 GB, own MCP server, run base 22000), Devstral on the Mac mini (run base 21000). Same A4H. Both ran in parallel, so durations are not comparable. Raw data: runs/stage1/baseline_qwen.json, runs/stage1/baseline_devstral.json.

Columns: Qwen / Devstral.

Task Cat Score End reason Tool calls Repair rate Hidden tests Qwen error messages (unique, shortened)
T01 A 48.3 / 0 loop / tool_budget 12 / 60 0.4 (2/5) / 0.8 4/12 / 0/0 [WRITE] An error occured during the save operation. The changes were not stored.
G0105 A 0 / 0 tool_budget / loop 60 / 9 – (0/0) / 0.5 0/0 / 0/0
G0017 B 0 / 0 loop / tool_budget 8 / 60 0.333 (1/3) / 1.0 0/0 / 0/0 Activation was cancelled.; Unexpected character 1..n; DDLS Z0GZ600H_I_PROJ_EFFORT was not activated
G0125 C 0 / 0 report / loop 7 / 8 – (0/0) / 0.0 0/0 / 0/0
G0128 D 0 / 0 loop / report 28 / 0 0.6 (3/5) / – 0/0 / 0/0 [WRITE] The statement CLASS ... IMPLEMENTATION. is unexpected; [WRITE] The statement CONSTRUCTOR is unexpected
G0139 E 0 / 0 loop / loop 12 / 7 0.75 (6/8) / 0.0 0/0 / 0/0 The formal parameter "LV_NIGHTLY" does not exist. However, the following parameters have similar names: "IV_NI
G0151 F 0 / 0 tool_budget / loop 60 / 18 1.0 (1/1) / 0.667 0/0 / 0/0 The syntax for a method specification is "objref->method" or "class=>method".; "VAL =" expected after "ROUND("
G0157 G 51.0 / 0 tool_budget / tool_budget 60 / 60 0.947 (18/19) / 0.964 4/10 / 0/0 The statement "50" is invalid. Check the spelling.; The statement "00" is invalid. Check the spelling.; The st
G0174 H 0 / 0 loop / report 8 / 5 – (0/0) / 0.0 0/0 / 0/0 Field "Z0GZC04U_FLSALE-FLOWER_TYPE" is unknown.
G0167 I 0 / 75.0 loop / loop 8 / 8 0.333 (1/3) / 0.333 0/0 / 11/11 The statement concluding with "...250" ended unexpectedly.; The statement "00" is invalid. Check the spelling.
G0185 K 75.0 / 0 loop / loop 20 / 14 0.7 (7/10) / 1.0 11/11 / 0/0 "VAL =" was expected, not "EV =".; "DEC =" was expected, not "DECIMALS =".; A class already exists with the na

Summary

Qwen 3.8 27B Devstral Small 2
Mean score 15.8 6.8
Tasks above 0 3/11 (T01 48.3, G0157 51.0, G0185 75.0) 1/11 (G0167 75.0)
End reasons loop 7, report 1, tool_budget 3 loop 6, report 2, tool_budget 3
Repair rate (pushes changed after an error / pushes after an error) 39/54 = 0.72 64/76 = 0.84

Findings

  • Qwen is better than Devstral on this subset (mean 15.8 against 6.8; 3 tasks above 0 against 1), but both are far from usable. Qwen has no empty responses with thinking off.
  • The tasks Qwen passes (T01, G0157, G0185) are different from Devstral's (G0167). Only 11 tasks: the difference is not statistically reliable.
  • Qwen loops or exhausts the budget in 10 of 11 tasks. G0157: 22 failed writes, 18 changed pushes after errors, still 51 points: repair works there. G0105 and G0151 hit the budget of 60 calls with only 0 and 2 failed writes; the reason is not analysed yet (the trajectories are in the run directories).
  • G0125 ended with report after 7 calls and no active object. G0167, where Devstral reached 75, ended with loop and score 0 for Qwen.
  • The reference for stage 1 training is this Qwen baseline: mean 15.8, 3 of 11 tasks above 0.

Not comparable: older Qwen runs in runs/archive/qwen38/ (thinking on, other limits).