G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 06:15:08 +02:00
parent 0c0e07fe96
commit c2c3803b1b
10 changed files with 470 additions and 14 deletions

View File

@@ -1,4 +1,4 @@
# EPOD: stale lock after a write (2 cases, not reproduced under control)
# EPOD: stale lock after a write (3 cases; the third one gives the cause)
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
@@ -13,6 +13,24 @@ Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide tha
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
## Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running
At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the
enqueue table (read with the reader class below, 2026-10-06 06:15):
```
GNAME=SEOCLSENQ GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====... GMODE=X GOBJ=ESEOCLASS GCLIENT=001 GUNAME=KESELI
GUSR=20261005194600195642000500vhcala4hci_A4H_00... GUSE=1 GTHOST=vhcala4hci_A4H_00 GTWP=05 GTDATE=20261005 GTTIME=194600 (server time, 2 hours behind CEST: 21:46)
```
Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning.
Its object was not cleaned up because the model had named it `ZCL_<run prefix>_JOB_COST_TEST` (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline
restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: `[LOCK] ... User KESELI is currently editing` on every `sap_push_element` (the run looped and scored 75).
Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.
So the cause is probably: **a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table** (the stateful ADT session that owns it is gone, and nothing removes the lock).
This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the *client* does not stop the server's write.
It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
@@ -53,8 +71,13 @@ There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `EN
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
## Also found: leftovers with the prefix inside the name
27 classes named `ZCL_<prefix>_...` / `ZCX_<prefix>_...` were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in
`sap_inactive_objects` and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, `harness/proxy.py`, `harness/runner.py`, `harness/sweep.py`).
## Wish for EPOD
0. **Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts** (a lock of user KESELI with the server's work process that has no live ADT session behind it).
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.