Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -18,7 +18,7 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
|
||||
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
|
||||
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
||||
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
||||
- Stage 1 reference baseline: Qwen, same settings, run on the MacBook (`baseline_qwen.json`, run base 22000).
|
||||
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
|
||||
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
|
||||
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
|
||||
|
||||
@@ -72,7 +72,7 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
|
||||
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
|
||||
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
|
||||
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
|
||||
- Result: mean 6.8; 1 of 11 tasks above 0 (G0167: 75). End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
|
||||
- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
|
||||
(the model changes the source, but the changes do not remove the cause).
|
||||
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
||||
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).
|
||||
|
||||
Reference in New Issue
Block a user