diff --git a/docs/epod-lock-leak.md b/docs/epod-lock-leak.md new file mode 100644 index 0000000..eb99dc0 --- /dev/null +++ b/docs/epod-lock-leak.md @@ -0,0 +1,22 @@ +# EPOD: stale lock after a write (one case, not reproduced) + +2026-10-05 15:09, run 200102 (task G1034, class `Z4AEE0SQ_STORAGE_PRICER`, two trajectory workers and three generators +running on A4H, one MCP session each): + +1. `sap_push_source` main, `sap_push_source includeType=testclasses` (twice, the second with two activation warnings), + `sap_check_object`, `sap_run_unit_test` all returned success. +2. `sap_push_element` then failed: `[LOCK] The object ... could not be locked. HTTP 403, ExceptionResourceNoAccess. + Server message: User KESELI is currently editing Z4AEE0SQ_STORAGE_PRICER.` Twice, minutes apart. +3. Teardown (ADT deletion API, after the MCP session was closed) failed: `You are already editing ...`. The lock stayed + until it was deleted in SM12 (2026-10-05 evening). + +So a lock was left behind by an earlier call of the same run and it did not end with the MCP session. The other +worker's `sap_run_unit_test` started 0.64 s after the second testclasses write (another session's call during a write). + +Not reproduced (A4H, probe objects, 2026-10-05): 25 write cycles alone; 60 write cycles overlapping 2690 unit test calls of a +second session; 40 cycles of testclasses write + `sap_check_object` (ATC, unit tests, coverage) + `sap_object_members` + +second testclasses write + `push_element`, with a second session running `sap_check_object`; `push_element` on a method of +the test class (returns "lives in the testclasses include", no lock). Frequency: 1 in about 90 trajectory runs. + +Wish for EPOD: release the lock in a `finally` path of every write tool, also after a failed or warned activation, and +release all locks of a session when the MCP session ends. diff --git a/train/STATE.md b/train/STATE.md index 8fd5c4d..9f242e1 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -159,3 +159,9 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training. - Tokens of accepted samples (20 tool schemas kept): p50 20147, p90 39458, p95 47694, max 71791, n 50. - Syntax hints (proxy syntaxCheck added): 1 in 1 runs. - Harness events: 2 ({'text:[LOCK]': 1, 'teardown': 1}); trajectory workers now 1. + +### Notes after the Opus review of the first stage 2 summary (2026-10-05 18:30) + +- **Token length (for the bf16 memory test):** accepted samples, 20 tool schemas kept: p50 20k, p90 39k, **p95 48k, max 72k** (about 10k tokens are the tool schemas). The bf16 memory test must use a **sequence length of 48k** (covers 95 % of the samples; samples over 48k are cut or dropped; 72k is not needed in the test). +- **DDLS reject reasons** (11 CDS runs: 6 accepted, 5 rejected): tool budget 2 (old limit 60, before the 100 budget), empty response 3 (one turn reached the 32k output limit while the model wrote its own CDS test class). Activation 0, hidden tests 0, ATC 0: all 11 runs passed every gate and every hidden test. Failed CDS tasks get the second attempt (pipeline rule). +- **G1034 lock event (cause not found):** after `sap_push_source includeType=testclasses` the class stayed locked ("User KESELI is currently editing", SM12 entry needed; teardown said "You are already editing"). Not a harness bug: the lock outlived the MCP session and the run. Not reproduced in 3 tests on A4H (alone, 60 writes overlapping 2690 unit test calls of another session, writes with ATC + coverage + member listing + a second session; `push_element` on a test-class method). One case in about 90 runs. Details for the EPOD developer: `docs/epod-lock-leak.md`. The harness treats it as an infrastructure event (run not accepted, worker count 2 to 1, 3 equal events stop the pipeline) and lists the object in the dashboard.