G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 06:15:08 +02:00
parent 0c0e07fe96
commit c2c3803b1b
10 changed files with 470 additions and 14 deletions

View File

@@ -180,3 +180,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- 2026-10-05 22:35 **pipeline stopped at 21:43 by an infrastructure outage, found 22:26 (my miss).** The MCP server (127.0.0.1:3000) refused connections for a short time; three runs got `URLError(ConnectionRefusedError)` and the pipeline counted three equal harness errors and stopped (rule). Not a model or harness bug: A4H was up (11 h), the MCP server answers again. New `harness/infra.py`: an outage of MCP or A4H is waited out (check every 30 s, two answers in a row), the run's objects are deleted, the same run starts again; it is not counted as an event or a failure (`runs/traj/outages.jsonl`). Same for series A (applies after a restart of `harness.localqwen`; the running series keeps the old code and would skip the task). The 3 interrupted runs were cleaned (11 objects) and get their second attempt. Budget: panel 50.84 at ledger 113.36 (2.54 ledger per usage in the last stretch); guard recomputed: panel left 9.16 - 3.0 reserve = 6.2 usage x 1.2 = 7.4 ledger, `BUDGET_LIMIT_USD` 129 (guard 121). Dashboard: a stopped pipeline shows no burn rate or ETA.
- 2026-10-05 22:40 cause of the 21:43 outage confirmed by Kral: he restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment; their objects could be deleted afterwards (no stale lock), so a restart of the MCP server in the middle of a write did not leak a lock in this case. Rule: restarting Eclipse is fine now (the pipeline waits and reruns), but check `python3 scripts_probe/lockprobe.py` for stale locks afterwards.
- 2026-10-06 06:40 **correction of the 22:40 entry and a new finding.** (1) An Eclipse restart in the middle of a write does leak a lock: `ZCL_Z4CGT1HJ_JOB_COST_TEST` (enqueue SEOCLSENQ, mode X, created 21:46, three minutes after the 21:43 restart) stayed locked until the next morning; details in `docs/epod-lock-leak.md` (case 3). I had missed it because the teardown only cleaned names that start with the run prefix. SM12 is needed (Kral). (2) **Leftover objects with the prefix inside the name** (ZCL_<prefix>_..., the model's own test classes): 27 stayed in A4H, and **57 of 90 accepted trajectories contain other runs' object names in tool results** (mostly `sap_inactive_objects`, in 30 cases the model pulled another run's class). Fixed: `harness/proxy.py` hides them (search, usage, inactive list) and answers "not available" for reads and writes of another run's object; `harness/runner.py` finds mid-name objects for teardown; `harness/sweep.py` deleted the leftovers (28 objects; the locked class remains). The 90 trajectories are not changed yet: see the data scrub option of `train/build_stage2.py`.