Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 17:01:52 +02:00
parent dbe6077a48
commit a368233d26
2 changed files with 41 additions and 2 deletions

39
docs/stage1-baseline.md Normal file
View File

@@ -0,0 +1,39 @@
# Stage 1 baseline: Qwen 3.8 27B vs Devstral Small 2
Date: 2026-10-04. Subset: `train/subset.json` (11 tasks). Same settings for both: no thinking (Qwen: `enable_thinking=false` per request; Devstral has no thinking mode), temperature 0.2, max_tokens 16384, tool-call budget 60, loop guard 3. Both models: MLX affine 4 bit (`Qwen3.8-27B-4bit`, `Devstral-Small-2-24B-4bit`).
Qwen ran on the MacBook (M4 Pro 48 GB, own MCP server, run base 22000), Devstral on the Mac mini (run base 21000). Same A4H. Both ran in parallel, so durations are not comparable. Raw data: `runs/stage1/baseline_qwen.json`, `runs/stage1/baseline_devstral.json`.
Columns: Qwen / Devstral.
| Task | Cat | Score | End reason | Tool calls | Repair rate | Hidden tests | Qwen error messages (unique, shortened) |
|---|---|---|---|---|---|---|---|
| T01 | A | **48.3** / 0 | loop / tool_budget | 12 / 60 | 0.4 (2/5) / 0.8 | 4/12 / 0/0 | [WRITE] An error occured during the save operation. The changes were not stored. |
| G0105 | A | **0** / 0 | tool_budget / loop | 60 / 9 | – (0/0) / 0.5 | 0/0 / 0/0 | |
| G0017 | B | **0** / 0 | loop / tool_budget | 8 / 60 | 0.333 (1/3) / 1.0 | 0/0 / 0/0 | Activation was cancelled.; Unexpected character 1..n; DDLS Z0GZ600H_I_PROJ_EFFORT was not activated |
| G0125 | C | **0** / 0 | report / loop | 7 / 8 | – (0/0) / 0.0 | 0/0 / 0/0 | |
| G0128 | D | **0** / 0 | loop / report | 28 / 0 | 0.6 (3/5) / – | 0/0 / 0/0 | [WRITE] The statement CLASS ... IMPLEMENTATION. is unexpected; [WRITE] The statement CONSTRUCTOR is unexpected |
| G0139 | E | **0** / 0 | loop / loop | 12 / 7 | 0.75 (6/8) / 0.0 | 0/0 / 0/0 | The formal parameter "LV_NIGHTLY" does not exist. However, the following parameters have similar names: "IV_NI |
| G0151 | F | **0** / 0 | tool_budget / loop | 60 / 18 | 1.0 (1/1) / 0.667 | 0/0 / 0/0 | The syntax for a method specification is "objref->method" or "class=>method".; "VAL =" expected after "ROUND(" |
| G0157 | G | **51.0** / 0 | tool_budget / tool_budget | 60 / 60 | 0.947 (18/19) / 0.964 | 4/10 / 0/0 | The statement "50" is invalid. Check the spelling.; The statement "00" is invalid. Check the spelling.; The st |
| G0174 | H | **0** / 0 | loop / report | 8 / 5 | – (0/0) / 0.0 | 0/0 / 0/0 | Field "Z0GZC04U_FLSALE-FLOWER_TYPE" is unknown. |
| G0167 | I | **0** / 75.0 | loop / loop | 8 / 8 | 0.333 (1/3) / 0.333 | 0/0 / 11/11 | The statement concluding with "...250" ended unexpectedly.; The statement "00" is invalid. Check the spelling. |
| G0185 | K | **75.0** / 0 | loop / loop | 20 / 14 | 0.7 (7/10) / 1.0 | 11/11 / 0/0 | "VAL =" was expected, not "EV =".; "DEC =" was expected, not "DECIMALS =".; A class already exists with the na |
## Summary
| | Qwen 3.8 27B | Devstral Small 2 |
|---|---|---|
| Mean score | **15.8** | 6.8 |
| Tasks above 0 | **3/11** (T01 48.3, G0157 51.0, G0185 75.0) | 1/11 (G0167 75.0) |
| End reasons | loop 7, report 1, tool_budget 3 | loop 6, report 2, tool_budget 3 |
| Repair rate (pushes changed after an error / pushes after an error) | 39/54 = 0.72 | 64/76 = 0.84 |
## Findings
- Qwen is better than Devstral on this subset (mean 15.8 against 6.8; 3 tasks above 0 against 1), but both are far from usable. Qwen has no empty responses with thinking off.
- The tasks Qwen passes (T01, G0157, G0185) are different from Devstral's (G0167). Only 11 tasks: the difference is not statistically reliable.
- Qwen loops or exhausts the budget in 10 of 11 tasks. G0157: 22 failed writes, 18 changed pushes after errors, still 51 points: repair works there. G0105 and G0151 hit the budget of 60 calls with only 0 and 2 failed writes; the reason is not analysed yet (the trajectories are in the run directories).
- G0125 ended with `report` after 7 calls and no active object. G0167, where Devstral reached 75, ended with `loop` and score 0 for Qwen.
- The reference for stage 1 training is this Qwen baseline: mean 15.8, 3 of 11 tasks above 0.
Not comparable: older Qwen runs in `runs/archive/qwen38/` (thinking on, other limits).

View File

@@ -18,7 +18,7 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further. was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), - Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1. empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000). - Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or `train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64). restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
@@ -72,7 +72,7 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). `train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`. `runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
- Result: mean 6.8; 1 of 11 tasks above 0 (G0167: 75). End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 - Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
(the model changes the source, but the changes do not remove the cause). (the model changes the source, but the changes do not remove the cause).
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`. - Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`). - The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).