Files
abap-llm/docs/epod-lock-leak.md

7.6 KiB

EPOD: stale lock after a write (3 cases; the third one gives the cause)

Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).

What happened (the two real cases)

Case 1 Case 2
When 15:09, run 200102, task G1034 19:25, run 200273, task G1091
Object class Z4AEE0SQ_STORAGE_PRICER table Z4AJ50UB_PO_HEAD (seed table of the task)
Situation normal run; the second sap_push_source includeType=testclasses returned success (two activation warnings); the next sap_push_element failed [LOCK] ... User KESELI is currently editing, twice; teardown through ADT deletion said You are already editing I started a second controller by mistake (same run number, same prefix, same seed table) and then killed both controller processes while the seed was installed. Teardown/deletion of the table said You are already editing
Lock gone only after SM12 (Kral) only after SM12 (Kral)

In both cases the lock outlived the MCP session and the run (hours). ENQUEUE_READ (see below) showed nothing after the SM12 delete.

Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running

At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the enqueue table (read with the reader class below, 2026-10-06 06:15):

GNAME=SEOCLSENQ  GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====...   GMODE=X  GOBJ=ESEOCLASS  GCLIENT=001  GUNAME=KESELI
GUSR=20261005194600195642000500vhcala4hci_A4H_00...   GUSE=1  GTHOST=vhcala4hci_A4H_00  GTWP=05  GTDATE=20261005  GTTIME=194600 (server time, 2 hours behind CEST: 21:46)

Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning. Its object was not cleaned up because the model had named it ZCL_<run prefix>_JOB_COST_TEST (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: [LOCK] ... User KESELI is currently editing on every sap_push_element (the run looped and scored 75). Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.

So the cause is probably: a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table (the stateful ADT session that owns it is gone, and nothing removes the lock). This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the client does not stop the server's write. It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.

What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)

Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it. None left a lock. Scripts: scripts_probe/ (killtest.py, killtabl.py, killtwo.py, killcreate.py, deltest.py, lockprobe.py).

  1. Client killed (SIGKILL) 0.03 to 0.8 s after sap_push_source includeType=testclasses was sent (the write needs about 0.7 s): 9 kills, no lock. The server finishes the write.
  2. Client killed during sap_push_source TABL (DDL, activation): 6 kills at 0.5 to 4.5 s (the write takes about 0.7 s): no lock.
  3. Two clients write the same table at the same time and both are killed: 8 delays (0.05 to 0.7 s): no lock.
  4. Two clients both sap_create_object + sap_push_source the same table and both are killed (Case 2 as exactly as possible): 8 delays (0.05 to 1.0 s): no lock.
  5. ADT deletion of the object while a write on it runs: the deletion is refused (You are already editing) or wins the race, no lock stays (6 delays).
  6. Earlier: 25 write cycles alone; 60 cycles overlapping 2690 sap_run_unit_test calls of another session; 40 cycles of testclasses write + sap_check_object (ATC, unit tests, coverage) + sap_object_members + second testclasses write + sap_push_element with a second session on sap_check_object; sap_push_element on a method of the test class (error text only, no lock).

So neither "client killed during a write" nor "two writers on one object" nor "delete during a write" leaks a lock on A4H in a short test. The two real cases may need a long running or slow call (the real runs were under load from 3 generators and 1 to 2 trajectory workers, calls then take seconds), or a state of the ADT stateful session that the probes did not reach.

How to see the enqueue state (exact method)

A class with IF_OO_ADT_CLASSRUN calls ENQUEUE_READ and is run with sap_run_class (class ZPROBE0EQ_LOCKS, in $TMP; source in scripts_probe/lockprobe.py):

CALL FUNCTION 'ENQUEUE_READ'
  EXPORTING gclient = sy-mandt gname = '' garg = '' guname = ''
  IMPORTING subrc = lv_subrc
  TABLES enq = lt_enq            "TYPE STANDARD TABLE OF seqg3
  EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.

A normal lock of a running write looks like this (seen in the 0.6 s test while a pipeline run was writing): GNAME=SEOCLSENQ GARG=Z4AJW0UK_CL_BERTH_FEE_TEST====... (lock object of the class include, owner user KESELI). After every probe the table was empty. A stale lock would show up as an entry that stays after the MCP session is closed: read the table, close the session, read again.

There is no function module on A4H to delete an enqueue entry (TFDIR: only ENQUEUE_READ, no ENQUEUE_DELETE): SM12 is the only way to release a stale lock, so every reproduction costs one SM12 delete.

What the harness does about it today

  • A write result with [LOCK] ("currently editing") is an infrastructure event: the run is not accepted, the trajectory workers go from 2 to 1, three equal events stop the pipeline.
  • A failed teardown (You are already editing) is listed as a leftover (dashboard); the object needs an SM12 delete, then the harness deletes it.
  • One controller only (flock on runs/pipeline/controller.lock), so two controllers cannot install the same run number twice any more (this was Case 2).
  • Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.

Also found: leftovers with the prefix inside the name

27 classes named ZCL_<prefix>_... / ZCX_<prefix>_... were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in sap_inactive_objects and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, harness/proxy.py, harness/runner.py, harness/sweep.py).

Wish for EPOD

  1. Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts (a lock of user KESELI with the server's work process that has no live ADT session behind it).
  2. Release the lock in a finally path of every write tool (also after a failed or warned activation).
  3. Release all locks of an MCP session when the session ends (DELETE /mcp) and when the connection breaks.
  4. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
  5. A short idle timeout for the stateful ADT session (the lock lasted hours).