85 lines
7.6 KiB
Markdown
85 lines
7.6 KiB
Markdown
# EPOD: stale lock after a write (3 cases; the third one gives the cause)
|
|
|
|
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
|
|
|
|
## What happened (the two real cases)
|
|
|
|
| | Case 1 | Case 2 |
|
|
|---|---|---|
|
|
| When | 15:09, run 200102, task G1034 | 19:25, run 200273, task G1091 |
|
|
| Object | class `Z4AEE0SQ_STORAGE_PRICER` | table `Z4AJ50UB_PO_HEAD` (seed table of the task) |
|
|
| Situation | normal run; the second `sap_push_source includeType=testclasses` returned success (two activation warnings); the next `sap_push_element` failed `[LOCK] ... User KESELI is currently editing`, twice; teardown through ADT deletion said `You are already editing` | I started a second controller by mistake (same run number, same prefix, same seed table) and then killed both controller processes while the seed was installed. Teardown/deletion of the table said `You are already editing` |
|
|
| Lock gone | only after SM12 (Kral) | only after SM12 (Kral) |
|
|
|
|
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
|
|
|
|
## Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running
|
|
|
|
At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the
|
|
enqueue table (read with the reader class below, 2026-10-06 06:15):
|
|
|
|
```
|
|
GNAME=SEOCLSENQ GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====... GMODE=X GOBJ=ESEOCLASS GCLIENT=001 GUNAME=KESELI
|
|
GUSR=20261005194600195642000500vhcala4hci_A4H_00... GUSE=1 GTHOST=vhcala4hci_A4H_00 GTWP=05 GTDATE=20261005 GTTIME=194600 (server time, 2 hours behind CEST: 21:46)
|
|
```
|
|
Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning.
|
|
Its object was not cleaned up because the model had named it `ZCL_<run prefix>_JOB_COST_TEST` (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline
|
|
restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: `[LOCK] ... User KESELI is currently editing` on every `sap_push_element` (the run looped and scored 75).
|
|
Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.
|
|
|
|
So the cause is probably: **a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table** (the stateful ADT session that owns it is gone, and nothing removes the lock).
|
|
This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the *client* does not stop the server's write.
|
|
It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.
|
|
|
|
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
|
|
|
|
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
|
|
None left a lock. Scripts: `scripts_probe/` (`killtest.py`, `killtabl.py`, `killtwo.py`, `killcreate.py`, `deltest.py`, `lockprobe.py`).
|
|
|
|
1. Client killed (`SIGKILL`) 0.03 to 0.8 s after `sap_push_source includeType=testclasses` was sent (the write needs about 0.7 s): 9 kills, no lock. The server finishes the write.
|
|
2. Client killed during `sap_push_source TABL` (DDL, activation): 6 kills at 0.5 to 4.5 s (the write takes about 0.7 s): no lock.
|
|
3. Two clients write the same table at the same time and both are killed: 8 delays (0.05 to 0.7 s): no lock.
|
|
4. Two clients both `sap_create_object` + `sap_push_source` the same table and both are killed (Case 2 as exactly as possible): 8 delays (0.05 to 1.0 s): no lock.
|
|
5. ADT deletion of the object while a write on it runs: the deletion is refused (`You are already editing`) or wins the race, no lock stays (6 delays).
|
|
6. Earlier: 25 write cycles alone; 60 cycles overlapping 2690 `sap_run_unit_test` calls of another session; 40 cycles of testclasses write + `sap_check_object` (ATC, unit tests, coverage) + `sap_object_members` + second testclasses write + `sap_push_element` with a second session on `sap_check_object`; `sap_push_element` on a method of the test class (error text only, no lock).
|
|
|
|
So neither "client killed during a write" nor "two writers on one object" nor "delete during a write" leaks a lock on A4H in a short test. The two real cases may need a long running or slow call (the real runs were under load from 3 generators and 1 to 2 trajectory workers, calls then take seconds), or a state of the ADT stateful session that the probes did not reach.
|
|
|
|
## How to see the enqueue state (exact method)
|
|
|
|
A class with `IF_OO_ADT_CLASSRUN` calls `ENQUEUE_READ` and is run with `sap_run_class` (class `ZPROBE0EQ_LOCKS`, in `$TMP`; source in `scripts_probe/lockprobe.py`):
|
|
|
|
```abap
|
|
CALL FUNCTION 'ENQUEUE_READ'
|
|
EXPORTING gclient = sy-mandt gname = '' garg = '' guname = ''
|
|
IMPORTING subrc = lv_subrc
|
|
TABLES enq = lt_enq "TYPE STANDARD TABLE OF seqg3
|
|
EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.
|
|
```
|
|
|
|
A normal lock of a running write looks like this (seen in the 0.6 s test while a pipeline run was writing):
|
|
`GNAME=SEOCLSENQ GARG=Z4AJW0UK_CL_BERTH_FEE_TEST====...` (lock object of the class include, owner user KESELI).
|
|
After every probe the table was empty. A stale lock would show up as an entry that stays after the MCP session is closed:
|
|
read the table, close the session, read again.
|
|
|
|
There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `ENQUEUE_READ`, no `ENQUEUE_DELETE`): SM12 is the only way to release a stale lock, so every reproduction costs one SM12 delete.
|
|
|
|
## What the harness does about it today
|
|
|
|
- A write result with `[LOCK]` ("currently editing") is an infrastructure event: the run is not accepted, the trajectory workers go from 2 to 1, three equal events stop the pipeline.
|
|
- A failed teardown (`You are already editing`) is listed as a leftover (dashboard); the object needs an SM12 delete, then the harness deletes it.
|
|
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
|
|
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
|
|
|
|
## Also found: leftovers with the prefix inside the name
|
|
27 classes named `ZCL_<prefix>_...` / `ZCX_<prefix>_...` were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in
|
|
`sap_inactive_objects` and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, `harness/proxy.py`, `harness/runner.py`, `harness/sweep.py`).
|
|
|
|
## Wish for EPOD
|
|
|
|
0. **Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts** (a lock of user KESELI with the server's work process that has no live ADT session behind it).
|
|
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
|
|
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
|
|
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
|
|
4. A short idle timeout for the stateful ADT session (the lock lasted hours).
|