Compare commits

..

2 Commits

Author SHA1 Message Date
Kral
1ddb9a7c65 Series A result: local Qwen 3 of 20 accepted, failures are loops and search loops; dashboard public copy
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 05:45:24 +02:00
Kral
01f3372e2c STATE: 21:43 outage was an Eclipse restart (confirmed)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 22:29:44 +02:00
3 changed files with 73 additions and 0 deletions

68
docs/local-qwen-run.md Normal file
View File

@@ -0,0 +1,68 @@
# Series A: local Qwen 3.8 27B on the new kinds (2026-10-05 22:24 to 2026-10-06 04:00)
Setup: `docs/remote-model.md`. Model served on the MacBook (`/Users/I301710/models/Qwen3.8-27B-4bit`, mlx-lm 0.32.0), harness, A4H and MCP on the Mac mini.
Settings as the official baseline: thinking off, max_tokens 16384, loop guard 3, tool budget 60, temperature 0.2, same system prompt.
Acceptance: score at least 80, end reason `report`. One task at a time. 20 tasks in 4.7 hours of run time (the 12 hour window was not used up);
no server pause, no infrastructure outage, no harness error.
## Result per kind
| kind | accepted / runs | mean score | failure types |
|---|---|---|---|
| INTF | 1/3 | 33.3 | loop 2, not_active 2, contract 2, hidden_tests_not_run 2 |
| TABL | 0/3 | 0.0 | contract 3, hidden_tests_not_run 3, loop 2, not_active 2, tool_budget 1 |
| STRU | 0/3 | 25.0 | loop 2, not_active 2, contract 2, hidden_tests_not_run 2, tool_budget 1 |
| MSAG | 2/3 | 66.7 | tool_budget 1, not_active 1, contract 1, hidden_tests_not_run 1 |
| EXC | 0/3 | 0.0 | loop 3, not_active 3, contract 3, hidden_tests_not_run 3 |
| DDLS | 0/5 | 31.0 | loop 3, not_active 3, contract 3, hidden_tests_not_run 3, tool_budget 1, low_score 1 |
Total: **3 of 20 accepted (15 %)**. Failure types over all runs: {'loop': 12, 'not_active': 13, 'contract': 14, 'hidden_tests_not_run': 14, 'tool_budget': 4, 'low_score': 1}.
## Reading
- **Loops are the failure.** 12 of the 17 failed runs ended with `end_reason loop` (9 times the same source pushed three times in a row, 3 times
the same read call with the same result). 10 of the 12 are short (1 to 6 minutes, 7 to 24 calls), two are long (59 and 71 minutes, 41 and 29 calls): Qwen pushes a source, gets an activation or save error
and pushes the same source again instead of repairing it. This is the weakness the baseline showed (loop 7 of 11) and the main target of the SFT.
Typical errors it did not repair: `The component C_ZONE has been declared multiple times`, `The statement 50 is unexpected`,
`Syntax error in <table>: DDL source could not be saved`, `FRIENDS ... expected after GLOBAL`, `Test method can be defined only in test classes`.
- Tool budget (60 calls), 4 runs, two different patterns the loop guard does not catch (the arguments differ each time): **three runs did nothing but `sap_search_object`**
(STRU G1947: 60 searches, MSAG G1940: 60, DDLS G1012: 56 searches, no write at all), and TABL G1901 pushed 49 different sources in a row (51 writes, never active).
The loop guard only sees identical calls; these are search and rewrite loops with changing arguments. A teacher trajectory that stops searching
and writes, and one that repairs after the first activation error, are what the SFT must show.
- Only INTF G1917 (10 calls) and MSAG G1902 and G1945 (13 and 19 calls) were solved cleanly: the small, regular tasks.
- Two runs reached 75 to 80 points but not an accepted state (STRU G1903 loop at the end, DDLS G1011 loop, DDLS G1031 low score after 56 calls and 64 minutes).
- Proxy syntax hint (local abaplint) was added in 12 calls (STRU G1903: 9, EXC G1124: 3). It did not help here: both runs still looped (the hint names a line and
a parser message, Qwen pushed the same source again).
- Speed: mean 14 minutes per run, median 5; the slow ones (STRU G1903 71 min, DDLS G1031 64 min, TABL G1910 59 min) are long
generations, not pauses.
- Compare with DeepSeek on the same kinds: see the pipeline summaries in `train/STATE.md` (teacher acceptance about 76 %).
## Per task
| task | kind | accepted | score | end | calls | minutes | failure types |
|---|---|---|---|---|---|---|---|
| G1900 | INTF | no | 0 | loop | 7 | 6 | loop, not_active, contract, hidden_tests_not_run |
| G1913 | INTF | no | 0 | loop | 8 | 5 | loop, not_active, contract, hidden_tests_not_run |
| G1917 | INTF | accepted | 100.0 | report | 10 | 7 | - |
| G1901 | TABL | no | 0 | tool_budget | 60 | 15 | tool_budget, contract, hidden_tests_not_run |
| G1910 | TABL | no | 0 | loop | 41 | 59 | loop, not_active, contract, hidden_tests_not_run |
| G1911 | TABL | no | 0 | loop | 15 | 2 | loop, not_active, contract, hidden_tests_not_run |
| G1903 | STRU | no | 75.0 | loop | 29 | 71 | loop |
| G1946 | STRU | no | 0 | loop | 24 | 4 | loop, not_active, contract, hidden_tests_not_run |
| G1947 | STRU | no | 0 | tool_budget | 60 | 5 | tool_budget, not_active, contract, hidden_tests_not_run |
| G1902 | MSAG | accepted | 100.0 | report | 19 | 9 | - |
| G1940 | MSAG | no | 0 | tool_budget | 60 | 6 | tool_budget, not_active, contract, hidden_tests_not_run |
| G1945 | MSAG | accepted | 100.0 | report | 13 | 3 | - |
| G1096 | EXC | no | 0 | loop | 14 | 3 | loop, not_active, contract, hidden_tests_not_run |
| G1104 | EXC | no | 0 | loop | 15 | 5 | loop, not_active, contract, hidden_tests_not_run |
| G1124 | EXC | no | 0 | loop | 8 | 1 | loop, not_active, contract, hidden_tests_not_run |
| G1000 | DDLS | no | 0 | loop | 8 | 2 | loop, not_active, contract, hidden_tests_not_run |
| G1011 | DDLS | no | 80.0 | loop | 18 | 4 | loop |
| G1012 | DDLS | no | 0 | tool_budget | 60 | 5 | tool_budget, not_active, contract, hidden_tests_not_run |
| G1020 | DDLS | no | 0 | loop | 11 | 3 | loop, not_active, contract, hidden_tests_not_run |
| G1031 | DDLS | no | 75.0 | report | 56 | 64 | low_score |
## Files
`runs/local_qwen/` (not in git): `results.jsonl`, `summary.json`, `state.json`, `preflight.json`, `runs/` (records), and `accepted.jsonl`
with the 3 accepted trajectories (`metadata.use_for_sft = false`, series `local_qwen_A`): not for the first SFT.

View File

@@ -459,6 +459,10 @@ def write_once():
tmp = OUT + ".tmp"
open(tmp, "w").write(page)
os.replace(tmp, OUT)
pub = os.path.join(OUT_DIR, "public") # the only folder that is served on the local network
os.makedirs(pub, exist_ok=True)
open(os.path.join(pub, "index.html.tmp"), "w").write(page)
os.replace(os.path.join(pub, "index.html.tmp"), os.path.join(pub, "index.html"))
last = d["hist"][-1]
with open(HISTORY, "a") as f:
f.write(json.dumps(last) + "\n")

View File

@@ -179,3 +179,4 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- 2026-10-05 21:31 budget correction 4: panel 50.00 at ledger 111.22; since 45.0 (100.62): 10.6 ledger / 5.0 usage = 2.12 (generation of new types and first runs). Guard recomputed with the pessimistic ratio 1.2: panel left 60 - 50 - 3.0 reserve = 7.0 usage x 1.2 = 8.4 ledger, guard 119.6, `BUDGET_LIMIT_USD` 128 (reserve 8). Only kinds below target run (Kral + Opus 2026-10-05).
- 2026-10-05 22:35 **pipeline stopped at 21:43 by an infrastructure outage, found 22:26 (my miss).** The MCP server (127.0.0.1:3000) refused connections for a short time; three runs got `URLError(ConnectionRefusedError)` and the pipeline counted three equal harness errors and stopped (rule). Not a model or harness bug: A4H was up (11 h), the MCP server answers again. New `harness/infra.py`: an outage of MCP or A4H is waited out (check every 30 s, two answers in a row), the run's objects are deleted, the same run starts again; it is not counted as an event or a failure (`runs/traj/outages.jsonl`). Same for series A (applies after a restart of `harness.localqwen`; the running series keeps the old code and would skip the task). The 3 interrupted runs were cleaned (11 objects) and get their second attempt. Budget: panel 50.84 at ledger 113.36 (2.54 ledger per usage in the last stretch); guard recomputed: panel left 9.16 - 3.0 reserve = 6.2 usage x 1.2 = 7.4 ledger, `BUDGET_LIMIT_USD` 129 (guard 121). Dashboard: a stopped pipeline shows no burn rate or ETA.
- 2026-10-05 22:40 cause of the 21:43 outage confirmed by Kral: he restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment; their objects could be deleted afterwards (no stale lock), so a restart of the MCP server in the middle of a write did not leak a lock in this case. Rule: restarting Eclipse is fine now (the pipeline waits and reruns), but check `python3 scripts_probe/lockprobe.py` for stale locks afterwards.