5.0 KiB
EPOD: stale lock after a write (2 cases, not reproduced under control)
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
What happened (the two real cases)
| Case 1 | Case 2 | |
|---|---|---|
| When | 15:09, run 200102, task G1034 | 19:25, run 200273, task G1091 |
| Object | class Z4AEE0SQ_STORAGE_PRICER |
table Z4AJ50UB_PO_HEAD (seed table of the task) |
| Situation | normal run; the second sap_push_source includeType=testclasses returned success (two activation warnings); the next sap_push_element failed [LOCK] ... User KESELI is currently editing, twice; teardown through ADT deletion said You are already editing |
I started a second controller by mistake (same run number, same prefix, same seed table) and then killed both controller processes while the seed was installed. Teardown/deletion of the table said You are already editing |
| Lock gone | only after SM12 (Kral) | only after SM12 (Kral) |
In both cases the lock outlived the MCP session and the run (hours). ENQUEUE_READ (see below) showed nothing after the SM12 delete.
What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
None left a lock. Scripts: scripts_probe/ (killtest.py, killtabl.py, killtwo.py, killcreate.py, deltest.py, lockprobe.py).
- Client killed (
SIGKILL) 0.03 to 0.8 s aftersap_push_source includeType=testclasseswas sent (the write needs about 0.7 s): 9 kills, no lock. The server finishes the write. - Client killed during
sap_push_source TABL(DDL, activation): 6 kills at 0.5 to 4.5 s (the write takes about 0.7 s): no lock. - Two clients write the same table at the same time and both are killed: 8 delays (0.05 to 0.7 s): no lock.
- Two clients both
sap_create_object+sap_push_sourcethe same table and both are killed (Case 2 as exactly as possible): 8 delays (0.05 to 1.0 s): no lock. - ADT deletion of the object while a write on it runs: the deletion is refused (
You are already editing) or wins the race, no lock stays (6 delays). - Earlier: 25 write cycles alone; 60 cycles overlapping 2690
sap_run_unit_testcalls of another session; 40 cycles of testclasses write +sap_check_object(ATC, unit tests, coverage) +sap_object_members+ second testclasses write +sap_push_elementwith a second session onsap_check_object;sap_push_elementon a method of the test class (error text only, no lock).
So neither "client killed during a write" nor "two writers on one object" nor "delete during a write" leaks a lock on A4H in a short test. The two real cases may need a long running or slow call (the real runs were under load from 3 generators and 1 to 2 trajectory workers, calls then take seconds), or a state of the ADT stateful session that the probes did not reach.
How to see the enqueue state (exact method)
A class with IF_OO_ADT_CLASSRUN calls ENQUEUE_READ and is run with sap_run_class (class ZPROBE0EQ_LOCKS, in $TMP; source in scripts_probe/lockprobe.py):
CALL FUNCTION 'ENQUEUE_READ'
EXPORTING gclient = sy-mandt gname = '' garg = '' guname = ''
IMPORTING subrc = lv_subrc
TABLES enq = lt_enq "TYPE STANDARD TABLE OF seqg3
EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.
A normal lock of a running write looks like this (seen in the 0.6 s test while a pipeline run was writing):
GNAME=SEOCLSENQ GARG=Z4AJW0UK_CL_BERTH_FEE_TEST====... (lock object of the class include, owner user KESELI).
After every probe the table was empty. A stale lock would show up as an entry that stays after the MCP session is closed:
read the table, close the session, read again.
There is no function module on A4H to delete an enqueue entry (TFDIR: only ENQUEUE_READ, no ENQUEUE_DELETE): SM12 is the only way to release a stale lock, so every reproduction costs one SM12 delete.
What the harness does about it today
- A write result with
[LOCK]("currently editing") is an infrastructure event: the run is not accepted, the trajectory workers go from 2 to 1, three equal events stop the pipeline. - A failed teardown (
You are already editing) is listed as a leftover (dashboard); the object needs an SM12 delete, then the harness deletes it. - One controller only (flock on
runs/pipeline/controller.lock), so two controllers cannot install the same run number twice any more (this was Case 2). - Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
Wish for EPOD
- Release the lock in a
finallypath of every write tool (also after a failed or warned activation). - Release all locks of an MCP session when the session ends (
DELETE /mcp) and when the connection breaks. - Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
- A short idle timeout for the stateful ADT session (the lock lasted hours).