Series A result: local Qwen 3 of 20 accepted, failures are loops and search loops; dashboard public copy
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
68
docs/local-qwen-run.md
Normal file
68
docs/local-qwen-run.md
Normal file
@@ -0,0 +1,68 @@
|
||||
# Series A: local Qwen 3.8 27B on the new kinds (2026-10-05 22:24 to 2026-10-06 04:00)
|
||||
|
||||
Setup: `docs/remote-model.md`. Model served on the MacBook (`/Users/I301710/models/Qwen3.8-27B-4bit`, mlx-lm 0.32.0), harness, A4H and MCP on the Mac mini.
|
||||
Settings as the official baseline: thinking off, max_tokens 16384, loop guard 3, tool budget 60, temperature 0.2, same system prompt.
|
||||
Acceptance: score at least 80, end reason `report`. One task at a time. 20 tasks in 4.7 hours of run time (the 12 hour window was not used up);
|
||||
no server pause, no infrastructure outage, no harness error.
|
||||
|
||||
## Result per kind
|
||||
|
||||
| kind | accepted / runs | mean score | failure types |
|
||||
|---|---|---|---|
|
||||
| INTF | 1/3 | 33.3 | loop 2, not_active 2, contract 2, hidden_tests_not_run 2 |
|
||||
| TABL | 0/3 | 0.0 | contract 3, hidden_tests_not_run 3, loop 2, not_active 2, tool_budget 1 |
|
||||
| STRU | 0/3 | 25.0 | loop 2, not_active 2, contract 2, hidden_tests_not_run 2, tool_budget 1 |
|
||||
| MSAG | 2/3 | 66.7 | tool_budget 1, not_active 1, contract 1, hidden_tests_not_run 1 |
|
||||
| EXC | 0/3 | 0.0 | loop 3, not_active 3, contract 3, hidden_tests_not_run 3 |
|
||||
| DDLS | 0/5 | 31.0 | loop 3, not_active 3, contract 3, hidden_tests_not_run 3, tool_budget 1, low_score 1 |
|
||||
|
||||
Total: **3 of 20 accepted (15 %)**. Failure types over all runs: {'loop': 12, 'not_active': 13, 'contract': 14, 'hidden_tests_not_run': 14, 'tool_budget': 4, 'low_score': 1}.
|
||||
|
||||
## Reading
|
||||
|
||||
- **Loops are the failure.** 12 of the 17 failed runs ended with `end_reason loop` (9 times the same source pushed three times in a row, 3 times
|
||||
the same read call with the same result). 10 of the 12 are short (1 to 6 minutes, 7 to 24 calls), two are long (59 and 71 minutes, 41 and 29 calls): Qwen pushes a source, gets an activation or save error
|
||||
and pushes the same source again instead of repairing it. This is the weakness the baseline showed (loop 7 of 11) and the main target of the SFT.
|
||||
Typical errors it did not repair: `The component C_ZONE has been declared multiple times`, `The statement 50 is unexpected`,
|
||||
`Syntax error in <table>: DDL source could not be saved`, `FRIENDS ... expected after GLOBAL`, `Test method can be defined only in test classes`.
|
||||
- Tool budget (60 calls), 4 runs, two different patterns the loop guard does not catch (the arguments differ each time): **three runs did nothing but `sap_search_object`**
|
||||
(STRU G1947: 60 searches, MSAG G1940: 60, DDLS G1012: 56 searches, no write at all), and TABL G1901 pushed 49 different sources in a row (51 writes, never active).
|
||||
The loop guard only sees identical calls; these are search and rewrite loops with changing arguments. A teacher trajectory that stops searching
|
||||
and writes, and one that repairs after the first activation error, are what the SFT must show.
|
||||
- Only INTF G1917 (10 calls) and MSAG G1902 and G1945 (13 and 19 calls) were solved cleanly: the small, regular tasks.
|
||||
- Two runs reached 75 to 80 points but not an accepted state (STRU G1903 loop at the end, DDLS G1011 loop, DDLS G1031 low score after 56 calls and 64 minutes).
|
||||
- Proxy syntax hint (local abaplint) was added in 12 calls (STRU G1903: 9, EXC G1124: 3). It did not help here: both runs still looped (the hint names a line and
|
||||
a parser message, Qwen pushed the same source again).
|
||||
- Speed: mean 14 minutes per run, median 5; the slow ones (STRU G1903 71 min, DDLS G1031 64 min, TABL G1910 59 min) are long
|
||||
generations, not pauses.
|
||||
- Compare with DeepSeek on the same kinds: see the pipeline summaries in `train/STATE.md` (teacher acceptance about 76 %).
|
||||
|
||||
## Per task
|
||||
|
||||
| task | kind | accepted | score | end | calls | minutes | failure types |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| G1900 | INTF | no | 0 | loop | 7 | 6 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1913 | INTF | no | 0 | loop | 8 | 5 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1917 | INTF | accepted | 100.0 | report | 10 | 7 | - |
|
||||
| G1901 | TABL | no | 0 | tool_budget | 60 | 15 | tool_budget, contract, hidden_tests_not_run |
|
||||
| G1910 | TABL | no | 0 | loop | 41 | 59 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1911 | TABL | no | 0 | loop | 15 | 2 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1903 | STRU | no | 75.0 | loop | 29 | 71 | loop |
|
||||
| G1946 | STRU | no | 0 | loop | 24 | 4 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1947 | STRU | no | 0 | tool_budget | 60 | 5 | tool_budget, not_active, contract, hidden_tests_not_run |
|
||||
| G1902 | MSAG | accepted | 100.0 | report | 19 | 9 | - |
|
||||
| G1940 | MSAG | no | 0 | tool_budget | 60 | 6 | tool_budget, not_active, contract, hidden_tests_not_run |
|
||||
| G1945 | MSAG | accepted | 100.0 | report | 13 | 3 | - |
|
||||
| G1096 | EXC | no | 0 | loop | 14 | 3 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1104 | EXC | no | 0 | loop | 15 | 5 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1124 | EXC | no | 0 | loop | 8 | 1 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1000 | DDLS | no | 0 | loop | 8 | 2 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1011 | DDLS | no | 80.0 | loop | 18 | 4 | loop |
|
||||
| G1012 | DDLS | no | 0 | tool_budget | 60 | 5 | tool_budget, not_active, contract, hidden_tests_not_run |
|
||||
| G1020 | DDLS | no | 0 | loop | 11 | 3 | loop, not_active, contract, hidden_tests_not_run |
|
||||
| G1031 | DDLS | no | 75.0 | report | 56 | 64 | low_score |
|
||||
|
||||
## Files
|
||||
|
||||
`runs/local_qwen/` (not in git): `results.jsonl`, `summary.json`, `state.json`, `preflight.json`, `runs/` (records), and `accepted.jsonl`
|
||||
with the 3 accepted trajectories (`metadata.use_for_sft = false`, series `local_qwen_A`): not for the first SFT.
|
||||
@@ -459,6 +459,10 @@ def write_once():
|
||||
tmp = OUT + ".tmp"
|
||||
open(tmp, "w").write(page)
|
||||
os.replace(tmp, OUT)
|
||||
pub = os.path.join(OUT_DIR, "public") # the only folder that is served on the local network
|
||||
os.makedirs(pub, exist_ok=True)
|
||||
open(os.path.join(pub, "index.html.tmp"), "w").write(page)
|
||||
os.replace(os.path.join(pub, "index.html.tmp"), os.path.join(pub, "index.html"))
|
||||
last = d["hist"][-1]
|
||||
with open(HISTORY, "a") as f:
|
||||
f.write(json.dumps(last) + "\n")
|
||||
|
||||
Reference in New Issue
Block a user