Files
abap-llm/docs/epod-lock-leak.md

5.0 KiB

EPOD: stale lock after a write (2 cases, not reproduced under control)

Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).

What happened (the two real cases)

Case 1 Case 2
When 15:09, run 200102, task G1034 19:25, run 200273, task G1091
Object class Z4AEE0SQ_STORAGE_PRICER table Z4AJ50UB_PO_HEAD (seed table of the task)
Situation normal run; the second sap_push_source includeType=testclasses returned success (two activation warnings); the next sap_push_element failed [LOCK] ... User KESELI is currently editing, twice; teardown through ADT deletion said You are already editing I started a second controller by mistake (same run number, same prefix, same seed table) and then killed both controller processes while the seed was installed. Teardown/deletion of the table said You are already editing
Lock gone only after SM12 (Kral) only after SM12 (Kral)

In both cases the lock outlived the MCP session and the run (hours). ENQUEUE_READ (see below) showed nothing after the SM12 delete.

What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)

Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it. None left a lock. Scripts: scripts_probe/ (killtest.py, killtabl.py, killtwo.py, killcreate.py, deltest.py, lockprobe.py).

  1. Client killed (SIGKILL) 0.03 to 0.8 s after sap_push_source includeType=testclasses was sent (the write needs about 0.7 s): 9 kills, no lock. The server finishes the write.
  2. Client killed during sap_push_source TABL (DDL, activation): 6 kills at 0.5 to 4.5 s (the write takes about 0.7 s): no lock.
  3. Two clients write the same table at the same time and both are killed: 8 delays (0.05 to 0.7 s): no lock.
  4. Two clients both sap_create_object + sap_push_source the same table and both are killed (Case 2 as exactly as possible): 8 delays (0.05 to 1.0 s): no lock.
  5. ADT deletion of the object while a write on it runs: the deletion is refused (You are already editing) or wins the race, no lock stays (6 delays).
  6. Earlier: 25 write cycles alone; 60 cycles overlapping 2690 sap_run_unit_test calls of another session; 40 cycles of testclasses write + sap_check_object (ATC, unit tests, coverage) + sap_object_members + second testclasses write + sap_push_element with a second session on sap_check_object; sap_push_element on a method of the test class (error text only, no lock).

So neither "client killed during a write" nor "two writers on one object" nor "delete during a write" leaks a lock on A4H in a short test. The two real cases may need a long running or slow call (the real runs were under load from 3 generators and 1 to 2 trajectory workers, calls then take seconds), or a state of the ADT stateful session that the probes did not reach.

How to see the enqueue state (exact method)

A class with IF_OO_ADT_CLASSRUN calls ENQUEUE_READ and is run with sap_run_class (class ZPROBE0EQ_LOCKS, in $TMP; source in scripts_probe/lockprobe.py):

CALL FUNCTION 'ENQUEUE_READ'
  EXPORTING gclient = sy-mandt gname = '' garg = '' guname = ''
  IMPORTING subrc = lv_subrc
  TABLES enq = lt_enq            "TYPE STANDARD TABLE OF seqg3
  EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.

A normal lock of a running write looks like this (seen in the 0.6 s test while a pipeline run was writing): GNAME=SEOCLSENQ GARG=Z4AJW0UK_CL_BERTH_FEE_TEST====... (lock object of the class include, owner user KESELI). After every probe the table was empty. A stale lock would show up as an entry that stays after the MCP session is closed: read the table, close the session, read again.

There is no function module on A4H to delete an enqueue entry (TFDIR: only ENQUEUE_READ, no ENQUEUE_DELETE): SM12 is the only way to release a stale lock, so every reproduction costs one SM12 delete.

What the harness does about it today

  • A write result with [LOCK] ("currently editing") is an infrastructure event: the run is not accepted, the trajectory workers go from 2 to 1, three equal events stop the pipeline.
  • A failed teardown (You are already editing) is listed as a leftover (dashboard); the object needs an SM12 delete, then the harness deletes it.
  • One controller only (flock on runs/pipeline/controller.lock), so two controllers cannot install the same run number twice any more (this was Case 2).
  • Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.

Wish for EPOD

  1. Release the lock in a finally path of every write tool (also after a failed or warned activation).
  2. Release all locks of an MCP session when the session ends (DELETE /mcp) and when the connection breaks.
  3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
  4. A short idle timeout for the stateful ADT session (the lock lasted hours).