Files
abap-llm/docs/epod-lock-leak.md

85 lines
7.6 KiB
Markdown

# EPOD: stale lock after a write (3 cases; the third one gives the cause)
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
## What happened (the two real cases)
| | Case 1 | Case 2 |
|---|---|---|
| When | 15:09, run 200102, task G1034 | 19:25, run 200273, task G1091 |
| Object | class `Z4AEE0SQ_STORAGE_PRICER` | table `Z4AJ50UB_PO_HEAD` (seed table of the task) |
| Situation | normal run; the second `sap_push_source includeType=testclasses` returned success (two activation warnings); the next `sap_push_element` failed `[LOCK] ... User KESELI is currently editing`, twice; teardown through ADT deletion said `You are already editing` | I started a second controller by mistake (same run number, same prefix, same seed table) and then killed both controller processes while the seed was installed. Teardown/deletion of the table said `You are already editing` |
| Lock gone | only after SM12 (Kral) | only after SM12 (Kral) |
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
## Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running
At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the
enqueue table (read with the reader class below, 2026-10-06 06:15):
```
GNAME=SEOCLSENQ GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====... GMODE=X GOBJ=ESEOCLASS GCLIENT=001 GUNAME=KESELI
GUSR=20261005194600195642000500vhcala4hci_A4H_00... GUSE=1 GTHOST=vhcala4hci_A4H_00 GTWP=05 GTDATE=20261005 GTTIME=194600 (server time, 2 hours behind CEST: 21:46)
```
Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning.
Its object was not cleaned up because the model had named it `ZCL_<run prefix>_JOB_COST_TEST` (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline
restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: `[LOCK] ... User KESELI is currently editing` on every `sap_push_element` (the run looped and scored 75).
Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.
So the cause is probably: **a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table** (the stateful ADT session that owns it is gone, and nothing removes the lock).
This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the *client* does not stop the server's write.
It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
None left a lock. Scripts: `scripts_probe/` (`killtest.py`, `killtabl.py`, `killtwo.py`, `killcreate.py`, `deltest.py`, `lockprobe.py`).
1. Client killed (`SIGKILL`) 0.03 to 0.8 s after `sap_push_source includeType=testclasses` was sent (the write needs about 0.7 s): 9 kills, no lock. The server finishes the write.
2. Client killed during `sap_push_source TABL` (DDL, activation): 6 kills at 0.5 to 4.5 s (the write takes about 0.7 s): no lock.
3. Two clients write the same table at the same time and both are killed: 8 delays (0.05 to 0.7 s): no lock.
4. Two clients both `sap_create_object` + `sap_push_source` the same table and both are killed (Case 2 as exactly as possible): 8 delays (0.05 to 1.0 s): no lock.
5. ADT deletion of the object while a write on it runs: the deletion is refused (`You are already editing`) or wins the race, no lock stays (6 delays).
6. Earlier: 25 write cycles alone; 60 cycles overlapping 2690 `sap_run_unit_test` calls of another session; 40 cycles of testclasses write + `sap_check_object` (ATC, unit tests, coverage) + `sap_object_members` + second testclasses write + `sap_push_element` with a second session on `sap_check_object`; `sap_push_element` on a method of the test class (error text only, no lock).
So neither "client killed during a write" nor "two writers on one object" nor "delete during a write" leaks a lock on A4H in a short test. The two real cases may need a long running or slow call (the real runs were under load from 3 generators and 1 to 2 trajectory workers, calls then take seconds), or a state of the ADT stateful session that the probes did not reach.
## How to see the enqueue state (exact method)
A class with `IF_OO_ADT_CLASSRUN` calls `ENQUEUE_READ` and is run with `sap_run_class` (class `ZPROBE0EQ_LOCKS`, in `$TMP`; source in `scripts_probe/lockprobe.py`):
```abap
CALL FUNCTION 'ENQUEUE_READ'
EXPORTING gclient = sy-mandt gname = '' garg = '' guname = ''
IMPORTING subrc = lv_subrc
TABLES enq = lt_enq "TYPE STANDARD TABLE OF seqg3
EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.
```
A normal lock of a running write looks like this (seen in the 0.6 s test while a pipeline run was writing):
`GNAME=SEOCLSENQ GARG=Z4AJW0UK_CL_BERTH_FEE_TEST====...` (lock object of the class include, owner user KESELI).
After every probe the table was empty. A stale lock would show up as an entry that stays after the MCP session is closed:
read the table, close the session, read again.
There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `ENQUEUE_READ`, no `ENQUEUE_DELETE`): SM12 is the only way to release a stale lock, so every reproduction costs one SM12 delete.
## What the harness does about it today
- A write result with `[LOCK]` ("currently editing") is an infrastructure event: the run is not accepted, the trajectory workers go from 2 to 1, three equal events stop the pipeline.
- A failed teardown (`You are already editing`) is listed as a leftover (dashboard); the object needs an SM12 delete, then the harness deletes it.
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
## Also found: leftovers with the prefix inside the name
27 classes named `ZCL_<prefix>_...` / `ZCX_<prefix>_...` were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in
`sap_inactive_objects` and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, `harness/proxy.py`, `harness/runner.py`, `harness/sweep.py`).
## Wish for EPOD
0. **Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts** (a lock of user KESELI with the server's work process that has no live ADT session behind it).
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
4. A short idle timeout for the stateful ADT session (the lock lasted hours).