G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
# EPOD: stale lock after a write (2 cases, not reproduced under control)
|
||||
# EPOD: stale lock after a write (3 cases; the third one gives the cause)
|
||||
|
||||
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
|
||||
|
||||
@@ -13,6 +13,24 @@ Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide tha
|
||||
|
||||
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
|
||||
|
||||
## Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running
|
||||
|
||||
At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the
|
||||
enqueue table (read with the reader class below, 2026-10-06 06:15):
|
||||
|
||||
```
|
||||
GNAME=SEOCLSENQ GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====... GMODE=X GOBJ=ESEOCLASS GCLIENT=001 GUNAME=KESELI
|
||||
GUSR=20261005194600195642000500vhcala4hci_A4H_00... GUSE=1 GTHOST=vhcala4hci_A4H_00 GTWP=05 GTDATE=20261005 GTTIME=194600 (server time, 2 hours behind CEST: 21:46)
|
||||
```
|
||||
Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning.
|
||||
Its object was not cleaned up because the model had named it `ZCL_<run prefix>_JOB_COST_TEST` (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline
|
||||
restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: `[LOCK] ... User KESELI is currently editing` on every `sap_push_element` (the run looped and scored 75).
|
||||
Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.
|
||||
|
||||
So the cause is probably: **a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table** (the stateful ADT session that owns it is gone, and nothing removes the lock).
|
||||
This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the *client* does not stop the server's write.
|
||||
It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.
|
||||
|
||||
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
|
||||
|
||||
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
|
||||
@@ -53,8 +71,13 @@ There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `EN
|
||||
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
|
||||
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
|
||||
|
||||
## Also found: leftovers with the prefix inside the name
|
||||
27 classes named `ZCL_<prefix>_...` / `ZCX_<prefix>_...` were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in
|
||||
`sap_inactive_objects` and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, `harness/proxy.py`, `harness/runner.py`, `harness/sweep.py`).
|
||||
|
||||
## Wish for EPOD
|
||||
|
||||
0. **Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts** (a lock of user KESELI with the server's work process that has no live ADT session behind it).
|
||||
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
|
||||
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
|
||||
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
|
||||
|
||||
Reference in New Issue
Block a user