Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
39
docs/stage1-baseline.md
Normal file
39
docs/stage1-baseline.md
Normal file
@@ -0,0 +1,39 @@
|
||||
# Stage 1 baseline: Qwen 3.8 27B vs Devstral Small 2
|
||||
|
||||
Date: 2026-10-04. Subset: `train/subset.json` (11 tasks). Same settings for both: no thinking (Qwen: `enable_thinking=false` per request; Devstral has no thinking mode), temperature 0.2, max_tokens 16384, tool-call budget 60, loop guard 3. Both models: MLX affine 4 bit (`Qwen3.8-27B-4bit`, `Devstral-Small-2-24B-4bit`).
|
||||
Qwen ran on the MacBook (M4 Pro 48 GB, own MCP server, run base 22000), Devstral on the Mac mini (run base 21000). Same A4H. Both ran in parallel, so durations are not comparable. Raw data: `runs/stage1/baseline_qwen.json`, `runs/stage1/baseline_devstral.json`.
|
||||
|
||||
Columns: Qwen / Devstral.
|
||||
|
||||
| Task | Cat | Score | End reason | Tool calls | Repair rate | Hidden tests | Qwen error messages (unique, shortened) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| T01 | A | **48.3** / 0 | loop / tool_budget | 12 / 60 | 0.4 (2/5) / 0.8 | 4/12 / 0/0 | [WRITE] An error occured during the save operation. The changes were not stored. |
|
||||
| G0105 | A | **0** / 0 | tool_budget / loop | 60 / 9 | – (0/0) / 0.5 | 0/0 / 0/0 | |
|
||||
| G0017 | B | **0** / 0 | loop / tool_budget | 8 / 60 | 0.333 (1/3) / 1.0 | 0/0 / 0/0 | Activation was cancelled.; Unexpected character 1..n; DDLS Z0GZ600H_I_PROJ_EFFORT was not activated |
|
||||
| G0125 | C | **0** / 0 | report / loop | 7 / 8 | – (0/0) / 0.0 | 0/0 / 0/0 | |
|
||||
| G0128 | D | **0** / 0 | loop / report | 28 / 0 | 0.6 (3/5) / – | 0/0 / 0/0 | [WRITE] The statement CLASS ... IMPLEMENTATION. is unexpected; [WRITE] The statement CONSTRUCTOR is unexpected |
|
||||
| G0139 | E | **0** / 0 | loop / loop | 12 / 7 | 0.75 (6/8) / 0.0 | 0/0 / 0/0 | The formal parameter "LV_NIGHTLY" does not exist. However, the following parameters have similar names: "IV_NI |
|
||||
| G0151 | F | **0** / 0 | tool_budget / loop | 60 / 18 | 1.0 (1/1) / 0.667 | 0/0 / 0/0 | The syntax for a method specification is "objref->method" or "class=>method".; "VAL =" expected after "ROUND(" |
|
||||
| G0157 | G | **51.0** / 0 | tool_budget / tool_budget | 60 / 60 | 0.947 (18/19) / 0.964 | 4/10 / 0/0 | The statement "50" is invalid. Check the spelling.; The statement "00" is invalid. Check the spelling.; The st |
|
||||
| G0174 | H | **0** / 0 | loop / report | 8 / 5 | – (0/0) / 0.0 | 0/0 / 0/0 | Field "Z0GZC04U_FLSALE-FLOWER_TYPE" is unknown. |
|
||||
| G0167 | I | **0** / 75.0 | loop / loop | 8 / 8 | 0.333 (1/3) / 0.333 | 0/0 / 11/11 | The statement concluding with "...250" ended unexpectedly.; The statement "00" is invalid. Check the spelling. |
|
||||
| G0185 | K | **75.0** / 0 | loop / loop | 20 / 14 | 0.7 (7/10) / 1.0 | 11/11 / 0/0 | "VAL =" was expected, not "EV =".; "DEC =" was expected, not "DECIMALS =".; A class already exists with the na |
|
||||
|
||||
## Summary
|
||||
|
||||
| | Qwen 3.8 27B | Devstral Small 2 |
|
||||
|---|---|---|
|
||||
| Mean score | **15.8** | 6.8 |
|
||||
| Tasks above 0 | **3/11** (T01 48.3, G0157 51.0, G0185 75.0) | 1/11 (G0167 75.0) |
|
||||
| End reasons | loop 7, report 1, tool_budget 3 | loop 6, report 2, tool_budget 3 |
|
||||
| Repair rate (pushes changed after an error / pushes after an error) | 39/54 = 0.72 | 64/76 = 0.84 |
|
||||
|
||||
## Findings
|
||||
|
||||
- Qwen is better than Devstral on this subset (mean 15.8 against 6.8; 3 tasks above 0 against 1), but both are far from usable. Qwen has no empty responses with thinking off.
|
||||
- The tasks Qwen passes (T01, G0157, G0185) are different from Devstral's (G0167). Only 11 tasks: the difference is not statistically reliable.
|
||||
- Qwen loops or exhausts the budget in 10 of 11 tasks. G0157: 22 failed writes, 18 changed pushes after errors, still 51 points: repair works there. G0105 and G0151 hit the budget of 60 calls with only 0 and 2 failed writes; the reason is not analysed yet (the trajectories are in the run directories).
|
||||
- G0125 ended with `report` after 7 calls and no active object. G0167, where Devstral reached 75, ended with `loop` and score 0 for Qwen.
|
||||
- The reference for stage 1 training is this Qwen baseline: mean 15.8, 3 of 11 tasks above 0.
|
||||
|
||||
Not comparable: older Qwen runs in `runs/archive/qwen38/` (thinking on, other limits).
|
||||
Reference in New Issue
Block a user