Lock leak: controlled reproduction attempts and method (not reproduced), enqueue reader; restart plan for 12 October

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 19:47:16 +02:00
parent 4432ac896c
commit 59496cb08b
9 changed files with 510 additions and 17 deletions

View File

@@ -1,22 +1,61 @@
# EPOD: stale lock after a write (one case, not reproduced)
# EPOD: stale lock after a write (2 cases, not reproduced under control)
2026-10-05 15:09, run 200102 (task G1034, class `Z4AEE0SQ_STORAGE_PRICER`, two trajectory workers and three generators
running on A4H, one MCP session each):
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
1. `sap_push_source` main, `sap_push_source includeType=testclasses` (twice, the second with two activation warnings),
`sap_check_object`, `sap_run_unit_test` all returned success.
2. `sap_push_element` then failed: `[LOCK] The object ... could not be locked. HTTP 403, ExceptionResourceNoAccess.
Server message: User KESELI is currently editing Z4AEE0SQ_STORAGE_PRICER.` Twice, minutes apart.
3. Teardown (ADT deletion API, after the MCP session was closed) failed: `You are already editing ...`. The lock stayed
until it was deleted in SM12 (2026-10-05 evening).
## What happened (the two real cases)
So a lock was left behind by an earlier call of the same run and it did not end with the MCP session. The other
worker's `sap_run_unit_test` started 0.64 s after the second testclasses write (another session's call during a write).
| | Case 1 | Case 2 |
|---|---|---|
| When | 15:09, run 200102, task G1034 | 19:25, run 200273, task G1091 |
| Object | class `Z4AEE0SQ_STORAGE_PRICER` | table `Z4AJ50UB_PO_HEAD` (seed table of the task) |
| Situation | normal run; the second `sap_push_source includeType=testclasses` returned success (two activation warnings); the next `sap_push_element` failed `[LOCK] ... User KESELI is currently editing`, twice; teardown through ADT deletion said `You are already editing` | I started a second controller by mistake (same run number, same prefix, same seed table) and then killed both controller processes while the seed was installed. Teardown/deletion of the table said `You are already editing` |
| Lock gone | only after SM12 (Kral) | only after SM12 (Kral) |
Not reproduced (A4H, probe objects, 2026-10-05): 25 write cycles alone; 60 write cycles overlapping 2690 unit test calls of a
second session; 40 cycles of testclasses write + `sap_check_object` (ATC, unit tests, coverage) + `sap_object_members` +
second testclasses write + `push_element`, with a second session running `sap_check_object`; `push_element` on a method of
the test class (returns "lives in the testclasses include", no lock). Frequency: 1 in about 90 trajectory runs.
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
Wish for EPOD: release the lock in a `finally` path of every write tool, also after a failed or warned activation, and
release all locks of a session when the MCP session ends.
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
None left a lock. Scripts: `scripts_probe/` (`killtest.py`, `killtabl.py`, `killtwo.py`, `killcreate.py`, `deltest.py`, `lockprobe.py`).
1. Client killed (`SIGKILL`) 0.03 to 0.8 s after `sap_push_source includeType=testclasses` was sent (the write needs about 0.7 s): 9 kills, no lock. The server finishes the write.
2. Client killed during `sap_push_source TABL` (DDL, activation): 6 kills at 0.5 to 4.5 s (the write takes about 0.7 s): no lock.
3. Two clients write the same table at the same time and both are killed: 8 delays (0.05 to 0.7 s): no lock.
4. Two clients both `sap_create_object` + `sap_push_source` the same table and both are killed (Case 2 as exactly as possible): 8 delays (0.05 to 1.0 s): no lock.
5. ADT deletion of the object while a write on it runs: the deletion is refused (`You are already editing`) or wins the race, no lock stays (6 delays).
6. Earlier: 25 write cycles alone; 60 cycles overlapping 2690 `sap_run_unit_test` calls of another session; 40 cycles of testclasses write + `sap_check_object` (ATC, unit tests, coverage) + `sap_object_members` + second testclasses write + `sap_push_element` with a second session on `sap_check_object`; `sap_push_element` on a method of the test class (error text only, no lock).
So neither "client killed during a write" nor "two writers on one object" nor "delete during a write" leaks a lock on A4H in a short test. The two real cases may need a long running or slow call (the real runs were under load from 3 generators and 1 to 2 trajectory workers, calls then take seconds), or a state of the ADT stateful session that the probes did not reach.
## How to see the enqueue state (exact method)
A class with `IF_OO_ADT_CLASSRUN` calls `ENQUEUE_READ` and is run with `sap_run_class` (class `ZPROBE0EQ_LOCKS`, in `$TMP`; source in `scripts_probe/lockprobe.py`):
```abap
CALL FUNCTION 'ENQUEUE_READ'
EXPORTING gclient = sy-mandt gname = '' garg = '' guname = ''
IMPORTING subrc = lv_subrc
TABLES enq = lt_enq "TYPE STANDARD TABLE OF seqg3
EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.
```
A normal lock of a running write looks like this (seen in the 0.6 s test while a pipeline run was writing):
`GNAME=SEOCLSENQ GARG=Z4AJW0UK_CL_BERTH_FEE_TEST====...` (lock object of the class include, owner user KESELI).
After every probe the table was empty. A stale lock would show up as an entry that stays after the MCP session is closed:
read the table, close the session, read again.
There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `ENQUEUE_READ`, no `ENQUEUE_DELETE`): SM12 is the only way to release a stale lock, so every reproduction costs one SM12 delete.
## What the harness does about it today
- A write result with `[LOCK]` ("currently editing") is an infrastructure event: the run is not accepted, the trajectory workers go from 2 to 1, three equal events stop the pipeline.
- A failed teardown (`You are already editing`) is listed as a leftover (dashboard); the object needs an SM12 delete, then the harness deletes it.
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
## Wish for EPOD
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.
4. A short idle timeout for the stateful ADT session (the lock lasted hours).

View File

@@ -0,0 +1,62 @@
# Restart plan for 12 October 2026 (cloud work)
Decision (Kral + Opus, 2026-10-05): when the budget guard stops, no cloud work (generation, trajectories) before the
Ollama reset on 12 October. Until then only no-cloud work. This is the plan for the restart.
## 1. Before the start (5 minutes, no cloud call)
1. Panel value after the reset (usage USD). Then:
`python3 -m harness.restart_plan --panel <value>`
It prints the `.env` values and the state below with live numbers. Put them in `.env`:
`BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD=<printed>`, `BUDGET_RESERVE_USD=8` (the 3 usage reserve for the
second teacher test stays). Ratio for the guard: 1.2 ledger per usage (trajectory runs; generation is about 2, so the
guard is on the safe side). Give the panel value again after about 2 hours and recompute (`python3 -m harness.dashboard panel <v>`).
2. A4H up (`docker ps`), MCP answers, no stale lock: `python3 scripts_probe/lockprobe.py` must print 0 entries.
3. One controller only: `python3 -m harness.pipeline` refuses to start a second one (`runs/pipeline/controller.lock`).
Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist.
4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`).
## 2. Order of work (what the controller does by itself)
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest
deficit goes first. Target (percent of accepted tasks and of accepted trajectories): CLAS 28, INTF 7, DDLS 25, FUNC 15,
PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5.
1. Trajectories for the new-type tasks that wait (INTF, TABL, STRU, MSAG, exception, PROG), then DDLS.
2. Second attempts: a task whose first attempt failed, or was accepted without a repair, gets one more attempt (at most two
attempts, at most two accepted trajectories per task). Failed DDLS and failed new-type tasks come first because they are
below target.
3. Generation (3 workers): the kind with the biggest deficit; a kind with more than 8 waiting tasks is not generated; a kind with
6 or more tries and under 20 % accepted is skipped. 20 % error-targeted slots, 30 % hard slots.
4. CLAS and FUNC generation and trajectories wait until they are at or below their target share (CLAS is far above).
5. K variants (free text) resume at 10 % of the other accepted tasks when their kinds are below target.
## 3. How much is needed (numbers of 2026-10-05 20:00, live: `restart_plan`)
CLAS has 49 accepted trajectories. At a 28 % share that is a total of 175 accepted trajectories, so about 105 more are
needed, all on other kinds: INTF 12, DDLS about 35, FUNC about 15, PROG about 17, TABL 14, STRU 3, MSAG 4, exception 4.
Tasks needed (accepted, with about 1.3 trajectory runs per accepted trajectory): TABL +9, INTF +10, STRU +3, MSAG +4,
DDLS +20, PROG +5. About 130 trajectory runs at 0.25 ledger = 33 ledger = 27 usage, plus the generation (about 6 usage).
Second attempts are part of the 130 runs. If the new reset gives 60 usage, this fits with room for the second teacher test.
## 4. Checks on the first runs of each new type (do not skip)
- The first 5 runs of INTF, TABL, STRU, MSAG, exception: records complete (the proxy must pass `sap_push_message`),
Qwen conversion works (`train/to_qwen.py` on `accepted.jsonl`: tool call round trip), no harness event.
- Acceptance per kind in the summary; a kind under 20 % after 6 tries is skipped automatically: read why (prompt, harness, or the
teacher cannot do it) before starting it again.
- DDLS: 8 of 17 runs accepted on 2026-10-05; the rejections were budget (60 calls, now 100) and empty responses (32k output limit).
If empty responses stay high, lower the output limit for CDS runs or retry once more.
## 5. Settings that stay
- Trajectory workers 2 (a harness event sets 1); stop rules: acceptance under 50 % over the last 30 runs, the same harness
error three times, the budget guard. Generation deadline 2026-10-10 18:00 has passed: set `--gen-deadline` and
`--traj-deadline` on the controller command (for example `--gen-deadline 2026-10-20T18:00 --traj-deadline 2026-10-21T23:30`).
- Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
- Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit.
## 6. After the restart
Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the
reserve), the EPOD requests in `docs/epod-lock-leak.md` and `docs/epod-syntax-hint.md`.