Compare commits

...

10 Commits

Author SHA1 Message Date
Kral
4261054eac Restart plan step 0b: rerun of the four empirical filter runs (G0183 G0180 G0143 G0185) without touching the eval set; empirical --rerun
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 09:07:10 +02:00
Kral
8ef85c2713 Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 09:04:09 +02:00
Kral
a465c33e3f D: own-test mutation scores of the accepted trajectories (metadata only), stage 2 set rebuilt with the score
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 07:36:39 +02:00
Kral
4002d889d5 own-test report and chain script; work list 2026-10-06 06:16:29 +02:00
Kral
c5f1a36df0 stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:16:02 +02:00
Kral
c2c3803b1b G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:15:08 +02:00
Kral
0c0e07fe96 F: eval slots for INTF/TABL/STRU/MSAG/exception (+K), Kral spot-check sheet, step 0 in the restart plan; D: own-test mutation scoring (running)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:06:41 +02:00
Kral
7f03849a86 C+E: stage 2 builder (mask, 48k, CLAS cap, family split, hook), private HF dataset, bf16 mixed training script with memory test, memory table, ratio proposal
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:01:49 +02:00
Kral
f5611cbf56 data analysis doc: corrected two counts 2026-10-06 05:56:31 +02:00
Kral
542fc8ec31 B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 05:56:23 +02:00
34 changed files with 2392 additions and 23 deletions

57
docs/bf16-memory.md Normal file
View File

@@ -0,0 +1,57 @@
# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
## Model (from config.json of Qwen3.8-27B)
64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**.
## Memory by component (GB; estimate, not a measurement)
| component | 48k | 64k | how |
|---|---|---|---|
| weights bf16 | 51.8 | 51.8 | 27.8 B x 2 bytes |
| LoRA weights + grads + 8-bit Adam | 1.0 | 1.0 | 107 M x 10 bytes |
| layer inputs for the backward pass | 29.3 on the GPU, about 0 with the Unsloth CPU offload (then 29 GB host RAM) | 39.1 / about 0 (host RAM 39 GB) | seq x 5120 x 2 bytes x 64 layers |
| recompute peak of one layer | 9.2 | 11.3 | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
| logits and loss | 3 chunked, 67 full logits | 3 chunked, 89 full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
## Fit by GPU (92 % of the card counted as usable)
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (model parallel) |
|---|---|---|---|---|---|---|---|
| 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits |
| 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits |
| 16k | on the GPU | chunked | 71 | fits | fits | fits | fits |
| 16k | on the GPU | full logits | 90 | no | tight | fits | fits |
| 32k | CPU offload (Unsloth) | chunked | 63 | fits | fits | fits | fits |
| 32k | CPU offload (Unsloth) | full logits | 104 | no | no | fits | fits |
| 32k | on the GPU | chunked | 82 | no | fits | fits | fits |
| 32k | on the GPU | full logits | 124 | no | no | fits | fits |
| 48k | CPU offload (Unsloth) | chunked | 65 | fits | fits | fits | fits |
| 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits |
| 48k | on the GPU | chunked | 94 | no | tight | fits | fits |
| 48k | on the GPU | full logits | 158 | no | no | no | fits |
| 64k | CPU offload (Unsloth) | chunked | 67 | fits | fits | fits | fits |
| 64k | CPU offload (Unsloth) | full logits | 153 | no | no | no | fits |
| 64k | on the GPU | chunked | 106 | no | no | fits | fits |
| 64k | on the GPU | full logits | 192 | no | no | no | fits |
Reading:
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about 65 GB at 48k, 67 GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
## Order when the test is run
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
A100 about 3.1 h (about 8 USD), H200 about 1.4 h (about 7 USD).
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
H200 about 4.3 h (about 21 USD), A100 about 9 h (about 23 USD). The real number comes from the memory test (`step_seconds`).

View File

@@ -0,0 +1,55 @@
# Stage 2 data analysis (accepted DeepSeek trajectories), 2026-10-06
Script `train/analyze.py`, numbers in `runs/analysis/analysis.json`. Base: 90 accepted trajectories of 119 runs (acceptance 76 %).
## 1. Repair taxonomy: are the three step-3 errors covered?
133 failed writes occurred inside accepted trajectories; each of the three common errors appears and is repaired:
| error (roadmap step 3) | failed writes | trajectories | repaired in the same trajectory | example |
|---|---|---|---|---|
| `TYPE c LENGTH n` in a method signature | 10 | 3 | 3 | Unable to interpret "4". Possible causes of error include incorrect spellings or comma errors. |
| reserved word as a field or parameter name | 18 | 16 | 16 | Field "VALUE" is unknown. |
| name longer than 30 characters | 25 | 23 | 23 | The name "SORTS_BY_TOTAL_DESC_THEN_VARIETY" is longer than the allowed 30 characters. |
Reading: the long names (23 trajectories) and the reserved words (16, by a text match: the label is generous, it also counts "Field VALUE is unknown") are well covered.
**`TYPE c LENGTH` in a signature is thin: 3 trajectories** (10 failed writes, all repaired). The SAP message is only "save operation failed" or, with the proxy hint,
`Statement does not exist ... METHODS`. This is the error that the proxy syntaxCheck exists for; after the reset the error-targeted slots ("named-type") should be
run first for CLAS and FUNC to get more of these. Other frequent errors (top list in the json): the testclasses include is not reachable with `sap_push_element` (9 times),
"save operation failed" without detail (30 with the hint), "Field X is unknown", "Test method can be defined only in test classes".
## 2. Behaviors that Qwen lacks (series A, all 20 runs) versus the teacher (accepted trajectories)
| behavior | DeepSeek (accepted) | Qwen series A |
|---|---|---|
| runs with at least one failed write | 60 of 90 | 14 of 20 |
| **repair after the first error** (a different source that is written successfully) | **59 of 60 (98 %)** | **6 of 14 (43 %)** |
| median calls from the first error to the repair | 1 | 2.0 |
| searches/reads before the first write (median / p90) | 3.0 / 19 | 4.0 / 15 |
| longest series of `sap_search_object` calls | 7 | **61** |
| runs with the same source pushed again | 2 of 90 | **11 of 20** |
| runs without any write | 4 | 4 |
The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61.
These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories).
Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection.
For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (29 runs), not analysed here.
## 3. Near duplicates
Pairwise comparison of the code that each accepted trajectory wrote (5-word shingles, prefix removed) and of the final reports: **no pair above 0.8 (code) or 0.85 (report)**, also none
between the two trajectories of one task. The data is not repetitive at this level.
## 4. empty_response (13 of 119 runs, 11 %)
By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and
contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call;
in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862,
the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 10 longer than 20000, 22 longer than 16000.
Fix: `docs/empty-response.md`.
## 5. What it means for the data
- Keep the repair behavior and the "write soon" behavior: they are present. Weight the error-targeted slots (named-type) up after the reset.
- Do not rely on `TYPE c LENGTH` repairs: only 3 examples.
- The DDLS share of empty runs (6 of 13) explains much of the CDS rejection.

23
docs/empty-response.md Normal file
View File

@@ -0,0 +1,23 @@
# empty_response fix (2026-10-06)
**Problem.** 13 of 119 DeepSeek runs (11 %) ended with `empty_response`: one turn used the whole output limit (32000 tokens) for reasoning,
with no content and no tool call, and the retry (same prompt, same temperature, up to two) did it again in 7 of 13 runs. Each such run costs up to
3 x 32000 output tokens and is lost (DDLS 6, PROG 3, CLAS 2, TABL 2). Analysis: `docs/data-analysis-stage2.md`, section 4.
**What the data says.** The empty turn comes after a small read result (50 to 6000 characters), in contexts of 13k to 63k tokens, not after a large result. Turns
over all runs: p50 1.1k, p90 7.9k tokens; legitimate long turns (large test classes) reach 30k; only 3 of 1997 turns longer than 24k were legitimate.
**Implemented (harness side only, `harness/agents.py`, `harness/trajectories.py`)**
1. Output cap per turn **24000** (was 32000; saves a quarter of every runaway turn, costs 3 of 1997 legitimate turns).
2. **One retry at temperature 0.8** (was two retries at 0.2): another sample instead of the same runaway.
3. **Stream guard** (off until verified): the request is streamed; when a turn has produced only reasoning for `STREAM_GUARD` tokens (suggested 9000) and no content and
no tool call, the stream is cut and the turn counts as empty (retry as in 2). Saves most of the cost of a runaway turn. Tested with a fake streaming server (text, tool call,
runaway: cut after 1500 estimated tokens in 0.01 s). **Not tested on the cloud model**: it needs the Ollama cloud stream to carry the reasoning in `delta.reasoning`
(or `reasoning_content` / `thinking`) and tool calls as streamed deltas; if the stream cannot be parsed the agent falls back to the normal request. Usage of a cut turn is estimated
(characters / 3.2) and marked `estimated` in the record.
4. Not implemented: "one object per write call" rule and splitting large test classes. Both change what the model sees (system prompt or task) and so the training distribution and
the comparison with the baseline (same system prompt). Option if 1 to 3 are not enough: a hint in the task spec of CDS tasks only.
**Verify after the reset (5 runs, DDLS and PROG first because they had most empty runs).** Run with `STREAM_GUARD=9000 python3 -m harness.pipeline` or set the constant:
check that `empty_response` stays below 5 % over 40 runs, that no accepted run was cut wrongly (`cut_by_stream_guard` in `metadata.turn_usage` followed by a good turn is fine),
and that the cost per run does not rise. If the stream breaks (tool calls missing), unset `STREAM_GUARD`: items 1 and 2 stay.

View File

@@ -1,4 +1,4 @@
# EPOD: stale lock after a write (2 cases, not reproduced under control) # EPOD: stale lock after a write (3 cases; the third one gives the cause)
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it). Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
@@ -13,6 +13,24 @@ Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide tha
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete. In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
## Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running
At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the
enqueue table (read with the reader class below, 2026-10-06 06:15):
```
GNAME=SEOCLSENQ GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====... GMODE=X GOBJ=ESEOCLASS GCLIENT=001 GUNAME=KESELI
GUSR=20261005194600195642000500vhcala4hci_A4H_00... GUSE=1 GTHOST=vhcala4hci_A4H_00 GTWP=05 GTDATE=20261005 GTTIME=194600 (server time, 2 hours behind CEST: 21:46)
```
Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning.
Its object was not cleaned up because the model had named it `ZCL_<run prefix>_JOB_COST_TEST` (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline
restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: `[LOCK] ... User KESELI is currently editing` on every `sap_push_element` (the run looped and scored 75).
Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.
So the cause is probably: **a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table** (the stateful ADT session that owns it is gone, and nothing removes the lock).
This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the *client* does not stop the server's write.
It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05) ## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it. Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
@@ -53,8 +71,13 @@ There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `EN
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2). - One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag. - Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
## Also found: leftovers with the prefix inside the name
27 classes named `ZCL_<prefix>_...` / `ZCX_<prefix>_...` were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in
`sap_inactive_objects` and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, `harness/proxy.py`, `harness/runner.py`, `harness/sweep.py`).
## Wish for EPOD ## Wish for EPOD
0. **Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts** (a lock of user KESELI with the server's work process that has no live ADT session behind it).
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation). 1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks. 2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session. 3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.

View File

@@ -0,0 +1,44 @@
# Eval seti: yeni türler için slot planı ve Kral'ın kontrol sayfası
Durum: hazırlık (2026-10-06). Üretim 12 Ekim sıfırlamasından sonra, **eğitim görevlerinden önce** (`docs/restart-12-oktober.md`, adım 0).
Amaç: eğitimden sonra INTF, TABL, STRU, MSAG ve istisna sınıfı görevlerini ölçebilmek. Şu an eval adaylarında (130 kabul) sadece CLAS, FUNC, DDLS, PROG var.
## 1. Slotlar (`python3 -m harness.evalset plan-new`)
Eğitim karışımıyla aynı paylar (110 görevlik sette INTF 8, TABL 9, STRU 2, MSAG 3, istisna 3); %30 fazla aday, çünkü ret ve ampirik süzgeç göreve göre eler.
| tür | kategori | slot | kimlikler | not |
|---|---|---|---|---|
| INTF | A | 10 | G0200–G0209 | arayüz: sabitler, tipler, metotlar; gizli test arayüzü uygulayan yerel sınıf kullanır |
| TABL | B | 11 | G0210–G0220 | tablo: alanlar, anahtar, not null; gizli test RTTI + `cl_osql_test_environment` |
| STRU | B | 3 | G0221–G0223 | yapı: alanlar, tipler; RTTI |
| MSAG | D | 4 | G0224–G0227 | mesaj sınıfı: numara, metin, yer tutucu; `MESSAGE ... INTO` |
| istisna (CLAS) | D | 4 | G0228–G0231 | `CX_` sınıfı: öznitelik, metin, zincir |
| K (serbest metin) | K | 5 | G0232–G0236 | INTF 1, TABL 2, MSAG 1, istisna 1; `evalset run-new-k`, tool isimleri EPOD'daki gibi (`generic_v0` kullanılmıyor, varsayım: bunu Kral onaylar) |
Her ikinci-üçüncü slot zor (zorluk 3). Her slot eğitim havuzuyla örtüşme kontrolünden geçer (eğitimdeki görevlere çok benziyorsa model yeniden yazar).
Maliyet tahmini: 32 slot x yaklaşık 0,15 ledger = 5 ledger, yaklaşık 2 usage; K varyantları 1 ledger'dan az.
## 2. Otomatik kapılar (Claude, senin işin değil)
oracle 100 ve null 0, mutasyon denetimi (INTF/TABL/STRU/MSAG/istisna için yeni mutantlar), eğitim havuzuyla örtüşme, ampirik süzgeç (DeepSeek ve yerel Qwen: seri A düzeneği eval görevlerini de koşturur),
sonra `docs/eval-inceleme.md` kontrol listesiyle Claude incelemesi (`python3 -m harness.review G0200`).
## 3. Senin kontrolün (yaklaşık 10 görev, yaklaşık 45 dakika)
Kimler: Claude'un **işaretlediği** görevler + her yeni türden **1 örnek** (5 görev) + 1 K varyantı. Her görev için `python3 -m harness.review <id>` tek ekranda spec, sözleşme, gizli test isimleri, mutasyon sonucu ve üretim geçmişini gösterir.
Kararın: `python3 -m harness.review <id> --set accept|fix|flag|reject --note "..."`.
Genel sorular (listedeki 1–11): spec belirsiz mi, her kuralın testi var mı, kontrat yeterli mi, zanaat serbest mi, gerçekçi mi, sızıntı var mı.
**Yeni türlere özel sorular (12–16):**
| # | Soru | tür |
|---|---|---|
| 12 | Spec her alanı, uzunluğu, anahtarı ve not-null'ı tek anlamlı veriyor mu? Gizli test **bayt/karakter** karışıklığına düşmüyor mu (RTTI `length` bayttır)? | TABL, STRU |
| 13 | Tablo testi gerçek davranışı ölçüyor mu (aynı anahtarla ikinci INSERT `sy-subrc 4`), yoksa sadece alan adlarına mı bakıyor? | TABL |
| 14 | Arayüz testi imza uyuşmazlığını yakalıyor mu (uygulayan yerel sınıf yalnızca imza doğruysa derlenir) ve sabit değerlerini okuyor mu? | INTF |
| 15 | Mesaj metinleri, numaralar ve yer tutucular spec'te **birebir** yazılı mı; test `MESSAGE ID ... INTO` ile son metni mi karşılaştırıyor? | MSAG |
| 16 | İstisna: öznitelik, metin ve (varsa) önceki istisna zinciri spec'te tanımlı mı; test fırlat-yakala ile ölçüyor mu? Gerçekten "istisna tasarımı" mı gerekiyor? | istisna |
## 4. Sonuç tablosu (üretimden sonra doldurulur)
| id | tür | otomatik kapılar | Claude incelemesi | Kral | not |
|---|---|---|---|---|---|
| | | | | | |

View File

@@ -0,0 +1,29 @@
# Other runs' objects in tool results: which runs are affected (2026-10-06)
Cause: the model sometimes names its own helper or test class `ZCL_<run prefix>_...` (prefix inside the name). The teardown looked only for names that start with the prefix, so such
classes stayed in A4H and showed up in `sap_inactive_objects`, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, `harness/sweep.py`).
Scan: `train/foreign_scan.py` (read only; result `runs/analysis/foreign_scan.json`). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object.
| group | runs | with foreign names in a tool result | with a foreign read | tools |
|---|---|---|---|---|
| official baseline Qwen (MacBook, 11 tasks) | 11 | 0 | 0 | - |
| Devstral baseline (dropped) | 11 | 0 | 0 | - |
| older Qwen baselines (archive) | 8 | 0 | 0 | - |
| smoke | 1 | 0 | 0 | - |
| eval empirical filter (DeepSeek on eval candidates) | 178 | 25 | 4 | sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5 |
| eval generation validation (oracle, null, mutants) | 1344 | 0 | 0 | - |
| early pilot runs | 16 | 1 | 0 | sap_inactive_objects 1 |
| training trajectories (DeepSeek) | 117 | 67 | 15 | sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1 |
| series A (local Qwen) | 20 | 3 | 0 | sap_inactive_objects 2, sap_search_object 1 |
## Reading
- **Official baseline (Qwen, 11 tasks, 2026-10-04): not affected** (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers,
and none of the 11 baseline runs called `sap_inactive_objects` (checked: 0 calls). **The baseline numbers stay as they are. Nothing was changed.**
- **Eval generation (1344 validation runs: oracle, null, mutants): not affected** (they call no list or search tool).
- **Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one**
(G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there
was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs).
- **Early pilot runs:** 1 of 16 saw a foreign name in an inactive list.
- **Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one.** Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool
(`train/requeue_foreign.py`, rows in `runs/traj/summary_excluded.jsonl`) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder.
- **Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads.** The scores (3 of 20) are not affected by a read; the series ran before the fix.

20
docs/own-test-mutation.md Normal file
View File

@@ -0,0 +1,20 @@
# Own-test mutation score of the accepted trajectories (item D, metadata only)
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
## Result
90 accepted trajectories looked at: {'scored': 63, 'not_supported': 5, 'no_own_tests': 22}.
- Scored: 63; the model's tests also passed on the correct reference: 53 (the others are not reliable: a test that fails on the correct solution kills every mutant).
- Mean score over the reliable ones: 0.99; median 1.00; distribution {'1.0': 50, '>=0.75': 3} (n = 53).
- By kind (mean, n): {CLAS: 1.00 (45), DDLS: 0.90 (5), FUNC: 1.00 (3)}
- `no_own_tests`: 22 trajectories were accepted without any own test (at most 85 points); `no_mutants`: 0; `not_supported` (PROG: tests are inside the program): 5.
## Weakest (score under 0.75, tests pass on the reference)
## Use
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).

View File

@@ -16,6 +16,24 @@ Ollama reset on 12 October. Until then only no-cloud work. This is the plan for
Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist. Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist.
4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`). 4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`).
## 1b. Step 0: eval tasks for the new kinds (before any training task)
`python3 -m harness.evalset plan-new` shows 32 slots (INTF 10, TABL 11, STRU 3, MSAG 4, exception 4; ids G0200 to G0231), then `python3 -m harness.evalset run-new`
(about 5 ledger, about 2 usage), then `python3 -m harness.evalset run-new-k` (5 K variants). Reason: the eval set (130 accepted candidates) has no INTF, TABL, STRU, MSAG
or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: `docs/eval-spotcheck-new-kinds.md`
(Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.
## 1c. Step 0b: rerun of four empirical filter runs (Kral + Opus 2026-10-06)
Four eval candidates saw or read leftover objects of other runs in the first empirical filter (`docs/foreign-objects-report.md`): **G0183, G0180, G0143, G0185** (DeepSeek, scores 80, 85, 85, 85).
After step 0 (eval generation for the new kinds), with the fixed proxy:
`python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 450000 --rerun G0183 G0180 G0143 G0185`
This writes only to `runs/emp_rerun/` (4 runs, about 1 ledger) and prints old and new filter decision per task. `empirical.json` and the eval set are **not** touched. Decision rules (step F):
easy candidate = score 95 or more and at most 20 tool calls; flag = score under 50 (a strong model fails although the reference passes); else normal. **If a decision changes for one of the four, report it to Kral and
Opus first; the eval set is changed only after their answer.** If nothing changes, say so in `docs/eval-inceleme.md` and leave the old results.
## 2. Order of work (what the controller does by itself) ## 2. Order of work (what the controller does by itself)
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest
@@ -56,6 +74,15 @@ Second attempts are part of the 130 runs. If the new reset gives 60 usage, this
- Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k. - Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
- Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit. - Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit.
## 5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks)
1. `git status` clean and pushed; `python3 -m harness.restart_plan --panel <value>` prints the numbers (panel value from Kral, expected 0 after the reset).
2. Services: A4H up, MCP answers, `python3 scripts_probe/lockprobe.py` shows no stale lock (only ATC runtime and debugger listener entries), `python3 -m harness.sweep` shows no leftover.
3. `.env`: `BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD`, `BUDGET_RESERVE_USD=8`; no `STOP` flag, no `STOPPED.txt`; no running controller (`runs/pipeline/controller.lock`).
4. Order of the first hour: step 0 eval slots for the new kinds (`evalset plan-new`, `run-new`, `run-new-k`), then the pending tasks of the kinds below target. The 9 trajectories that were
moved back (`runs/traj/summary_excluded.jsonl`: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order.
5. Settings to confirm: workers 2, `STREAM_GUARD` unset for the first runs and then 9000 for 5 DDLS/PROG runs (`docs/empty-response.md`), tool budget 100 for CDS.
6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host.
## 6. After the restart ## 6. After the restart
Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the

77
docs/stage2-build.md Normal file
View File

@@ -0,0 +1,77 @@
# Stage 2 training set, build of 2026-10-06 09:03:41
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
## Result
81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
Dropped:
- DDLS: {'over_48k': 5}
- INTF: {'over_48k': 1}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
## Not good enough yet (kinds under the minimum of 25)
| kind | samples | families | short |
|---|---|---|---|
| CLAS | 14 | 14 | 11 |
| INTF | 2 | 2 | 23 |
| DDLS | 9 | 9 | 16 |
| FUNC | 8 | 8 | 17 |
| PROG | 5 | 5 | 20 |
| TABL | 2 | 2 | 23 |
| STRU | 0 | 0 | 25 |
| MSAG | 0 | 0 | 25 |
| EXC | 0 | 0 | 25 |
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
| 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M |
| 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M |
| 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M |
| 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M |
| 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M |
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
| class of the trajectory | weight | reason |
|---|---|---|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
## Other decisions of 2026-10-06
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
- **Memory test:** waits until the data is near the size of the first SFT run.
## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).

22
epod_tests/README.md Normal file
View File

@@ -0,0 +1,22 @@
# EPOD acceptance tests (for the EPOD server developer)
Small test set for the two server changes that the harness work asked for (details: `docs/epod-syntax-hint.md`, `docs/epod-lock-leak.md`)
and for the parallel-call behavior (`harness/loadtest.py` numbers). Standard library only, Python 3.9+, no harness code needed.
It uses probe objects `ZEPODT_*` in `$TMP` on a **test system**.
```sh
cd epod_tests
export MCP_URL=http://127.0.0.1:3000/mcp MCP_TOKEN=... # the MCP server
export A4H_URL=http://localhost:50000 A4H_USER=... A4H_PASSWORD=... # only for the cleanup (ADT deletion); without it the tests list the probe objects
python3 run_tests.py # all; --only T1 T3 for some
```
| test | passes when | state of the server on 2026-10-06 |
|---|---|---|
| T1 syntax_hint | a rejected write (`TYPE c LENGTH 4` in a method signature) returns syntax messages with a line, not only "save operation failed" | FAIL expected (not implemented; the harness proxy adds abaplint messages) |
| T2 no_lock_after_kill | a client killed with SIGKILL during a write (4 delays) leaves no lock: the next write works | PASS (not reproducible on A4H) |
| T3 parallel_reads | 6 clients search + syntax check in parallel without `Concurrent call detected` | PASS |
| T4 parallel_writes | 3 clients create + write in parallel without `Concurrent call detected` | FAIL expected (the shared connection; 15 to 50 lock hits per 45 s in the harness load test) |
| T5 lock_error_text | informational | PASS |
A change in the server is done when T1 and T4 pass and T2 and T3 stay green. T1 accepts any answer that carries a line number or a `syntaxCheck` object with messages (the proxy format is in `docs/epod-syntax-hint.md`).

98
epod_tests/epod_client.py Normal file
View File

@@ -0,0 +1,98 @@
"""Minimal MCP client and ADT deletion for the EPOD acceptance tests (standard library only, Python 3.9+)."""
import base64
import http.cookiejar
import json
import os
import re
import time
import urllib.request
from xml.sax.saxutils import quoteattr
URL = os.environ.get("MCP_URL", "http://127.0.0.1:3000/mcp")
TOKEN = os.environ.get("MCP_TOKEN", "")
SYSTEM = os.environ.get("MCP_SYSTEM_ID") # optional: Eclipse project name when several systems are connected
class Mcp:
def __init__(self, timeout=300):
self.sid, self.n, self.timeout = None, 0, timeout
def _post(self, body, method="POST"):
h = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"}
if TOKEN:
h["Authorization"] = "Bearer " + TOKEN
if self.sid:
h["Mcp-Session-Id"] = self.sid
req = urllib.request.Request(URL, json.dumps(body).encode() if body is not None else None, h, method=method)
with urllib.request.urlopen(req, timeout=self.timeout) as r:
sid = r.headers.get("Mcp-Session-Id")
raw = r.read().decode()
self.sid = sid or self.sid
if "data:" in raw[:40]:
raw = "".join(l[5:].strip() for l in raw.splitlines() if l.startswith("data:"))
return json.loads(raw) if raw.strip() else None
def open(self):
self.n += 1
self._post({"jsonrpc": "2.0", "id": self.n, "method": "initialize", "params": {
"protocolVersion": "2025-03-26", "capabilities": {}, "clientInfo": {"name": "epod-acceptance", "version": "1"}}})
self._post({"jsonrpc": "2.0", "method": "notifications/initialized"})
return self
def close(self):
try:
self._post(None, "DELETE")
except Exception:
pass
def __enter__(self):
return self.open()
def __exit__(self, *a):
self.close()
def call(self, tool, args):
"""Returns (is_error, text). 'Concurrent call detected' is returned as it is (the tests count it)."""
self.n += 1
if SYSTEM:
args = dict(args, systemId=SYSTEM)
res = self._post({"jsonrpc": "2.0", "id": self.n, "method": "tools/call", "params": {"name": tool, "arguments": args}})
if "error" in res:
return True, json.dumps(res["error"])
r = res["result"]
return bool(r.get("isError")), "\n".join(c.get("text", "") for c in r.get("content", []))
class Adt:
"""Deletion through the ADT deletion API (needs A4H_URL, A4H_USER, A4H_PASSWORD; A4H_CLIENT default 001)."""
def __init__(self):
self.base = os.environ.get("A4H_URL", "").rstrip("/")
self.client = os.environ.get("A4H_CLIENT", "001")
self.auth = "Basic " + base64.b64encode(("%s:%s" % (os.environ.get("A4H_USER", ""), os.environ.get("A4H_PASSWORD", ""))).encode()).decode()
self.opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(http.cookiejar.CookieJar()))
self.csrf = None
def available(self):
return bool(self.base and os.environ.get("A4H_USER"))
def _req(self, path, body=None, headers=None):
h = {"Authorization": self.auth}
if self.csrf:
h["x-csrf-token"] = self.csrf
h.update(headers or {})
req = urllib.request.Request("%s%s%ssap-client=%s" % (self.base, path, "&" if "?" in path else "?", self.client), body.encode() if body else None, h)
with self.opener.open(req, timeout=120) as r:
return dict(r.headers), r.read().decode()
def delete(self, uris):
if not self.available() or not uris:
return {}
hdr, _ = self._req("/sap/bc/adt/discovery", headers={"x-csrf-token": "fetch", "Accept": "*/*"})
self.csrf = hdr.get("x-csrf-token") or hdr.get("X-CSRF-Token")
objs = "".join("<del:object adtcore:uri=%s><del:transportNumber/></del:object>" % quoteattr(u) for u in uris)
body = ('<?xml version="1.0" encoding="UTF-8"?><del:deletionRequest xmlns:del="http://www.sap.com/adt/deletion" '
'xmlns:adtcore="http://www.sap.com/adt/core">' + objs + "</del:deletionRequest>")
_, text = self._req("/sap/bc/adt/deletion/delete", body, {"Content-Type": "application/vnd.sap.adt.deletion.request.v1+xml",
"Accept": "application/vnd.sap.adt.deletion.response.v1+xml"})
return {m.group(1): 'isDeleted="true"' in m.group(0) for m in re.finditer(r'<del:object\b[^>]*adtcore:uri="([^"]+)"[^>]*>', text)}

View File

@@ -0,0 +1,32 @@
[
{
"id": "T1",
"name": "syntax_hint",
"pass": false,
"detail": "the failed write returned: {\"success\":false,\"error\":\"[WRITE] An error occured during the save operation. The changes were not stored.\"}"
},
{
"id": "T2",
"name": "no_lock_after_kill",
"pass": true,
"detail": "4 kills (0.05 to 0.7 s), the next write always worked"
},
{
"id": "T3",
"name": "parallel_reads",
"pass": true,
"detail": "372 rounds, 0 'Concurrent call detected'"
},
{
"id": "T4",
"name": "parallel_writes",
"pass": false,
"detail": "27 create+write rounds with 3 clients, 12 'Concurrent call detected'"
},
{
"id": "T5",
"name": "lock_error_text",
"pass": true,
"detail": "informational: see docs/epod-lock-leak.md (no way to create a lock on purpose)"
}
]

174
epod_tests/run_tests.py Normal file
View File

@@ -0,0 +1,174 @@
"""EPOD acceptance tests (2026-10-06). Run against an EPOD MCP server on a test system (probe objects ZEPODT_*, package $TMP).
MCP_URL=http://127.0.0.1:3000/mcp MCP_TOKEN=... [A4H_URL=http://localhost:50000 A4H_USER=... A4H_PASSWORD=...] python3 run_tests.py [--only T1 T2 ...]
Each test prints PASS or FAIL with the reason and writes the result to epod_tests_result.json. Tests:
T1 syntax_hint a rejected write ('save operation failed') comes with syntax messages (line and text) docs/epod-syntax-hint.md
T2 no_lock_after_kill a client killed during a write leaves no lock (the next write and the deletion work) docs/epod-lock-leak.md
T3 parallel_reads 6 clients read/ATC/syntax-check in parallel without 'Concurrent call detected'
T4 parallel_writes 3 clients create/write/activate in parallel without 'Concurrent call detected' (a fix: per system or per session lock)
T5 lock_error_text a write on a locked object names the lock owner (informational)
"""
import json
import os
import signal
import subprocess
import sys
import threading
import time
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from epod_client import Adt, Mcp # noqa: E402
CLS = """CLASS {n} DEFINITION PUBLIC FINAL CREATE PUBLIC.
PUBLIC SECTION.
METHODS run RETURNING VALUE(rv) TYPE i.
ENDCLASS.
CLASS {n} IMPLEMENTATION.
METHOD run.
rv = 1.
ENDMETHOD.
ENDCLASS.
"""
BAD = """CLASS {n} DEFINITION PUBLIC FINAL CREATE PUBLIC.
PUBLIC SECTION.
METHODS run IMPORTING iv_zone TYPE c LENGTH 4 RETURNING VALUE(rv) TYPE i.
ENDCLASS.
CLASS {n} IMPLEMENTATION.
METHOD run.
rv = 1.
ENDMETHOD.
ENDCLASS.
"""
TC = """CLASS ltc DEFINITION FINAL FOR TESTING DURATION SHORT RISK LEVEL HARMLESS.
PRIVATE SECTION.
METHODS t1 FOR TESTING.
ENDCLASS.
CLASS ltc IMPLEMENTATION.
METHOD t1.
cl_abap_unit_assert=>assert_equals( act = 1 exp = 1 ).
ENDMETHOD.
ENDCLASS.
"""
created = []
RUN = time.strftime("%H%M%S")
def new_class(m, tag):
name = "ZEPODT_%s_%s" % (tag, RUN)
m.call("sap_create_object", {"objectType": "CLAS", "objectName": name, "packageName": "$TMP", "description": "epod acceptance test"})
created.append("/sap/bc/adt/oo/classes/" + name.lower())
return name
def ok(text):
return '"success":true' in text.replace(" ", "")
def t1_syntax_hint():
with Mcp() as m:
n = new_class(m, "T1")
e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": n, "source": BAD.format(n=n.lower())})
low = t.lower()
has_detail = ("syntaxcheck" in low or "syntax" in low and "line" in low) and "save operation" in low or "line 3" in low or '"line":3' in low.replace(" ", "")
return bool(has_detail), "the failed write returned: " + t[:300]
def t2_no_lock_after_kill():
victim = ("import sys,json,os\nsys.path.insert(0,%r)\nfrom epod_client import Mcp\na=json.loads(sys.argv[1])\nm=Mcp().open()\nprint('SENT',flush=True)\nm.call('sap_push_source',a)\n"
% os.path.dirname(os.path.abspath(__file__)))
bad = []
for delay in (0.05, 0.2, 0.4, 0.7):
with Mcp() as m:
n = new_class(m, "T2")
m.call("sap_push_source", {"objectType": "CLAS", "objectName": n, "source": CLS.format(n=n.lower())})
p = subprocess.Popen([sys.executable, "-c", victim, json.dumps({"objectType": "CLAS", "objectName": n, "includeType": "testclasses", "source": TC})],
stdout=subprocess.PIPE, text=True)
p.stdout.readline()
time.sleep(delay)
try:
os.kill(p.pid, signal.SIGKILL)
except ProcessLookupError:
pass
p.wait()
time.sleep(3)
with Mcp() as m:
e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": n, "includeType": "testclasses", "source": TC + "* again\n"})
if not ok(t):
bad.append((delay, t[:160]))
return not bad, "after the kill the next write failed at: %s" % bad if bad else "4 kills (0.05 to 0.7 s), the next write always worked"
def parallel(n_clients, seconds, work):
out, end = [], time.time() + seconds
def run(i):
locks = calls = 0
with Mcp() as m:
while time.time() < end:
t = work(m, i)
calls += 1
locks += "Concurrent call detected" in t
out.append((calls, locks))
th = [threading.Thread(target=run, args=(i,)) for i in range(n_clients)]
[x.start() for x in th]
[x.join() for x in th]
return sum(c for c, _ in out), sum(l for _, l in out)
def t3_parallel_reads():
def work(m, i):
return m.call("sap_search_object", {"query": "CL_ABAP_CHAR_UTIL*", "objType": "CLAS"})[1] + \
m.call("sap_syntax_check", {"objectType": "CLAS", "objectName": "CL_ABAP_CHAR_UTILITIES"})[1]
calls, locks = parallel(6, 15, work)
return locks == 0, "%d rounds, %d 'Concurrent call detected'" % (calls, locks)
def t4_parallel_writes():
cnt = {"n": 0}
def work(m, i):
cnt["n"] += 1
name = "ZEPODT_P%d_%s%03d" % (i, RUN, cnt["n"] % 1000)
t = m.call("sap_create_object", {"objectType": "CLAS", "objectName": name, "packageName": "$TMP", "description": "parallel"})[1]
created.append("/sap/bc/adt/oo/classes/" + name.lower())
return t + m.call("sap_push_source", {"objectType": "CLAS", "objectName": name, "source": CLS.format(n=name.lower())})[1]
calls, locks = parallel(3, 25, work)
return locks == 0, "%d create+write rounds with 3 clients, %d 'Concurrent call detected'" % (calls, locks)
def t5_lock_error_text():
return True, "informational: see docs/epod-lock-leak.md (no way to create a lock on purpose)"
TESTS = [("T1", "syntax_hint", t1_syntax_hint), ("T2", "no_lock_after_kill", t2_no_lock_after_kill), ("T3", "parallel_reads", t3_parallel_reads),
("T4", "parallel_writes", t4_parallel_writes), ("T5", "lock_error_text", t5_lock_error_text)]
def main():
only = set(sys.argv[sys.argv.index("--only") + 1:]) if "--only" in sys.argv else None
result = []
for tid, name, fn in TESTS:
if only and tid not in only:
continue
t0 = time.time()
try:
passed, why = fn()
except Exception as e: # noqa: BLE001
passed, why = False, "error: %r" % e
print("%-4s %-22s %s (%.0f s) %s" % (tid, name, "PASS" if passed else "FAIL", time.time() - t0, why[:230]), flush=True)
result.append({"id": tid, "name": name, "pass": passed, "detail": why[:600]})
adt = Adt()
if adt.available() and created:
try:
gone = adt.delete(created)
print("cleanup: %d of %d probe objects deleted" % (sum(gone.values()), len(created)))
except Exception as e: # noqa: BLE001
print("cleanup failed:", repr(e)[:200])
elif created:
print("delete these probe objects yourself (SE80 or ADT): ZEPODT_* in $TMP (%d objects)" % len(created))
json.dump(result, open("epod_tests_result.json", "w"), indent=1)
if __name__ == "__main__":
main()

View File

@@ -84,7 +84,7 @@ class LlmAgent:
def __init__(self, model, base_url=None, api_key=None, max_turns=80, temperature=0.2, def __init__(self, model, base_url=None, api_key=None, max_turns=80, temperature=0.2,
max_seconds=None, max_tokens=None, chat_template_kwargs=None, loop_guard=None, max_seconds=None, max_tokens=None, chat_template_kwargs=None, loop_guard=None,
deadline=None, watch=False): deadline=None, watch=False, empty_retries=2, retry_temperature=None, stream_guard=None):
self.model = model self.model = model
self.name = f"llm:{model}" self.name = f"llm:{model}"
self.base_url = (base_url or os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1")).rstrip("/") self.base_url = (base_url or os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1")).rstrip("/")
@@ -96,7 +96,12 @@ class LlmAgent:
# Runaway reasoning: 35 of 1965 DeepSeek turns produced 393k output tokens and no content (46 % of the # Runaway reasoning: 35 of 1965 DeepSeek turns produced 393k output tokens and no content (46 % of the
# run cost, 2026-10-03). Normal turns: p95 17.5k, max 82k. A cut turn is retried (see run). # run cost, 2026-10-03). Normal turns: p95 17.5k, max 82k. A cut turn is retried (see run).
self.max_tokens = max_tokens or (32000 if ":cloud" in (model or "") else None) self.max_tokens = max_tokens or (32000 if ":cloud" in (model or "") else None)
self.empty_retries = 2 self.empty_retries = empty_retries
# empty_response fix (2026-10-06): a retry after an empty turn can use another temperature (the same sample often
# runs away again: 7 of 13 empty runs had two or three capped turns in a row); stream_guard = reasoning tokens
# after which a streamed turn without any content or tool call is cut and counted as an empty turn
self.retry_temperature = retry_temperature
self.stream_guard = stream_guard
self.loop_guard = loop_guard # end the run after this many identical pushes in a row (None = off) self.loop_guard = loop_guard # end the run after this many identical pushes in a row (None = off)
self.deadline = deadline # absolute time (time.time()) of the hard stop, or None self.deadline = deadline # absolute time (time.time()) of the hard stop, or None
self.watch = watch # remote model: ping the server during a request, end the run when it is gone self.watch = watch # remote model: ping the server during a request, end the run when it is gone
@@ -104,9 +109,9 @@ class LlmAgent:
self.messages, self.tools, self.reasoning, self.turn_usage = [], [], [], [] # for the trajectory record self.messages, self.tools, self.reasoning, self.turn_usage = [], [], [], [] # for the trajectory record
self.chat_template_kwargs = chat_template_kwargs # local server only, e.g. {"enable_thinking": False} self.chat_template_kwargs = chat_template_kwargs # local server only, e.g. {"enable_thinking": False}
def _chat(self, messages, tools): def _chat(self, messages, tools, temperature=None):
body = {"model": self.model, "messages": messages, "tools": tools, body = {"model": self.model, "messages": messages, "tools": tools,
"temperature": self.temperature, "parallel_tool_calls": False} "temperature": temperature if temperature is not None else self.temperature, "parallel_tool_calls": False}
if self.max_tokens: if self.max_tokens:
body["max_tokens"] = self.max_tokens body["max_tokens"] = self.max_tokens
if self.chat_template_kwargs: if self.chat_template_kwargs:
@@ -115,6 +120,11 @@ class LlmAgent:
{"Content-Type": "application/json", {"Content-Type": "application/json",
"Authorization": f"Bearer {self.api_key}"}) "Authorization": f"Bearer {self.api_key}"})
last = None last = None
if self.stream_guard and not (self.watch or self.deadline):
try:
return self._chat_stream(dict(body), messages)
except Exception as e: # noqa: BLE001 streaming not usable (server, parse): the normal request below
last = e
for attempt in range(4): # model server errors (HTTP 5xx, timeouts): retry with backoff for attempt in range(4): # model server errors (HTTP 5xx, timeouts): retry with backoff
try: try:
if self.watch or self.deadline: if self.watch or self.deadline:
@@ -137,6 +147,49 @@ class LlmAgent:
time.sleep(10 * (attempt + 1)) time.sleep(10 * (attempt + 1))
raise RuntimeError(f"model request failed after retries: {last}") raise RuntimeError(f"model request failed after retries: {last}")
def _chat_stream(self, body, messages):
"""Streamed request. A turn that has produced only reasoning for `stream_guard` tokens (about 3.2 characters per
token) and no content and no tool call is cut and returned as an empty turn (the run loop retries it)."""
body["stream"] = True
body["stream_options"] = {"include_usage": True}
req = urllib.request.Request(f"{self.base_url}/chat/completions", json.dumps(body).encode(),
{"Content-Type": "application/json", "Authorization": f"Bearer {self.api_key}"})
content, reasoning, calls, usage, cut = "", 0, {}, {}, False
with urllib.request.urlopen(req, timeout=self.request_timeout) as r:
for raw in r:
line = raw.decode("utf-8", "replace").strip()
if not line.startswith("data:"):
continue
data = line[5:].strip()
if data == "[DONE]":
break
chunk = json.loads(data)
if chunk.get("usage"):
usage = chunk["usage"]
for ch in chunk.get("choices") or []:
d = ch.get("delta") or {}
content += d.get("content") or ""
reasoning += len(d.get("reasoning") or d.get("reasoning_content") or d.get("thinking") or "")
for tc in d.get("tool_calls") or []:
c = calls.setdefault(tc.get("index", 0), {"id": tc.get("id") or "", "type": "function",
"function": {"name": "", "arguments": ""}})
c["id"] = c["id"] or tc.get("id") or ""
f = tc.get("function") or {}
c["function"]["name"] += f.get("name") or ""
a = f.get("arguments")
c["function"]["arguments"] += a if isinstance(a, str) else json.dumps(a) if a else ""
if not content.strip() and not calls and reasoning / 3.2 >= self.stream_guard:
cut = True
break
if cut:
est = {"prompt_tokens": int(len(json.dumps(messages)) / 3.5), "completion_tokens": int(reasoning / 3.2),
"estimated": True, "cut_by_stream_guard": True}
return {"role": "assistant", "content": ""}, est
msg = {"role": "assistant", "content": content or None}
if calls:
msg["tool_calls"] = [calls[k] for k in sorted(calls)]
return msg, usage
def _ping(self, timeout=10): def _ping(self, timeout=10):
try: try:
urllib.request.urlopen(urllib.request.Request(f"{self.base_url}/models"), timeout=timeout).read() urllib.request.urlopen(urllib.request.Request(f"{self.base_url}/models"), timeout=timeout).read()
@@ -195,7 +248,8 @@ class LlmAgent:
self.end_reason = "time_budget" self.end_reason = "time_budget"
break break
try: try:
msg, usage = self._chat(messages, tools) msg, usage = self._chat(messages, tools,
self.retry_temperature if (empty and self.retry_temperature) else None)
add_usage(self.model, usage, kind="run", ref=proxy.prefix) add_usage(self.model, usage, kind="run", ref=proxy.prefix)
self.turn_usage.append(usage) self.turn_usage.append(usage)
except WindowEnd: except WindowEnd:

View File

@@ -302,6 +302,22 @@ def local_card():
return hdr + tbl return hdr + tbl
def work_card():
"""Progress of the no-cloud work packages (runs/dashboard/work.json, updated by Claude after each item)."""
path = os.path.join(OUT_DIR, "work.json")
if not os.path.exists(path):
return ""
try:
w = json.load(open(path))
except ValueError:
return ""
cls = {"done": "ok", "in progress": "wa", "waiting": "gr", "parked": "gr", "blocked": "er"}
rows = "".join("<tr><td><b>%s</b></td><td>%s</td><td><span class='chip %s'>%s</span></td><td>%s</td></tr>" % (
E(i["id"]), E(i["text"]), cls.get(i["status"], "gr"), E(i["status"]), E(i.get("note", ""))) for i in w.get("items", []))
return ("<div class=card style='grid-column:1/-1'><h2>%s</h2><div class=b><table><tr><th>item</th><th>work</th><th>status</th><th>result</th></tr>%s</table></div>"
"<div class=note>updated %s</div></div>" % (E(w.get("title", "")), rows, time.strftime("%H:%M", time.localtime(w.get("updated", 0)))))
def mix_table(rows, evs): def mix_table(rows, evs):
kt = mix.accepted_task_counts() kt = mix.accepted_task_counts()
kr = {} kr = {}
@@ -437,7 +453,7 @@ def render(d):
% ("ok" if d["a4h"] else "er", "up" if d["a4h"] else "down", "ok" if d["mcp"] else "er", "up" if d["mcp"] else "down", scls, status, % ("ok" if d["a4h"] else "er", "up" if d["a4h"] else "down", "ok" if d["mcp"] else "er", "up" if d["mcp"] else "down", scls, status,
E("\n".join(d["procs"])))) E("\n".join(d["procs"]))))
logc = "<div class=card><h2>Pipeline log</h2><div class=b><div class='log mono'>%s</div></div></div>" % E("\n".join(d["log_tail"])) logc = "<div class=card><h2>Pipeline log</h2><div class=b><div class='log mono'>%s</div></div></div>" % E("\n".join(d["log_tail"]))
body = (banner + "<div class=tiles>" + tiles + "</div><div class=grid>" + local_card() + budget + prog + stops + sysc body = (banner + "<div class=tiles>" + tiles + "</div><div class=grid>" + work_card() + local_card() + budget + prog + stops + sysc
+ mix_table(rows, evs) + agg_table("Trajectories by category", evs, rows, "category") + agg_table("Trajectories by object type", evs, rows, "object_type") + mix_table(rows, evs) + agg_table("Trajectories by category", evs, rows, "category") + agg_table("Trajectories by object type", evs, rows, "object_type")
+ gen_table(logs, "category", "Task generation by category") + gen_table(logs, "object_type", "Task generation by object type") + gen_table(logs, "category", "Task generation by category") + gen_table(logs, "object_type", "Task generation by object type")
+ tok + recent + logc + "</div>") + tok + recent + logc + "</div>")

View File

@@ -40,6 +40,43 @@ def candidates():
return out return out
def decision(score, calls):
"""The filter decisions of step F: an easy candidate (DeepSeek >= 95 and <= 20 tool calls, tasks_gen/eval/easy_candidates.json),
a flag (a strong model fails, score under 50, although the reference passes: the spec may be unclear), else normal."""
if score is None:
return "no result"
if score >= 95 and (calls or 99) <= 20:
return "easy candidate"
return "flag: strong model fails" if score < 50 else "normal"
def rerun(a):
"""Rerun some tasks (after the proxy fix of 2026-10-06) WITHOUT touching empirical.json or the eval set: results go to runs/emp_rerun/,
then the old and the new filter decision are printed. A changed decision is reported to Kral and Opus before anything in the eval set changes."""
out_dir = os.path.join(ROOT, "runs", "emp_rerun")
os.makedirs(out_dir, exist_ok=True)
runner = Runner(POOL, out_dir)
changed = []
for i, tid in enumerate(a.tasks):
old = _json(os.path.join(POOL, tid, "empirical.json"), {}).get(a.model, {})
try:
rep, run_dir = runner.run(tid, LlmAgent(a.model, a.base_url), a.run_base + i)
except BudgetExceeded as e:
print("BUDGET", e, flush=True)
break
h = rep.get("hidden_tests") or {}
new = {"score": (rep.get("score") or {}).get("total"), "hidden": f"{h.get('passed')}/{h.get('total')}", "tool_calls": rep.get("tool_calls"),
"end_reason": rep.get("end_reason"), "run_dir": os.path.relpath(run_dir, ROOT)}
d_old, d_new = decision(old.get("score"), old.get("tool_calls")), decision(new["score"], new["tool_calls"])
line = {"task": tid, "old": {k: old.get(k) for k in ("score", "hidden", "tool_calls")}, "new": new, "decision_old": d_old, "decision_new": d_new,
"decision_changed": d_old != d_new}
open(os.path.join(out_dir, "results.jsonl"), "a").write(json.dumps(line) + "\n")
print(json.dumps(line), flush=True)
if line["decision_changed"]:
changed.append(tid)
print("DECISION CHANGED for:", changed or "none", "(nothing in tasks_gen/eval was changed)", flush=True)
def main(): def main():
load_env(os.path.join(ROOT, ".env")) load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser() ap = argparse.ArgumentParser()
@@ -47,7 +84,12 @@ def main():
ap.add_argument("--model", required=True) ap.add_argument("--model", required=True)
ap.add_argument("--base-url") ap.add_argument("--base-url")
ap.add_argument("--run-base", type=int, required=True) ap.add_argument("--run-base", type=int, required=True)
ap.add_argument("--rerun", action="store_true", help="rerun the named tasks into runs/emp_rerun/ and compare the filter decision; the eval set is not changed")
a = ap.parse_args() a = ap.parse_args()
if a.rerun:
if not a.tasks:
raise SystemExit("--rerun needs task ids")
return rerun(a)
runs_root = os.path.join(ROOT, "runs", "emp") runs_root = os.path.join(ROOT, "runs", "emp")
os.makedirs(runs_root, exist_ok=True) os.makedirs(runs_root, exist_ok=True)
runner = Runner(POOL, runs_root) runner = Runner(POOL, runs_root)

View File

@@ -16,6 +16,7 @@ import sys
from .adt_client import load_env from .adt_client import load_env
from .generator import ROOT, generate, make_k_variant from .generator import ROOT, generate, make_k_variant
from . import mix, overlap
from .ledger import BudgetExceeded, spent from .ledger import BudgetExceeded, spent
FIRST_ID = 100 FIRST_ID = 100
@@ -52,6 +53,73 @@ K_SOURCES = [("G0004", "free_text"), ("G0007", "incomplete"), ("G0026", "free_te
RELEASES = ["v702", "v740sp05"] RELEASES = ["v702", "v740sp05"]
# New kinds for the eval set (Opus item F, 2026-10-06): the 130 eval candidates have only CLAS, FUNC, DDLS and PROG, so INTF, TABL, STRU, MSAG and
# exception tasks cannot be measured after training. Shares as in the training mix (110 tasks: INTF 8, TABL 9, STRU 2, MSAG 3, exception 3); about
# 30 percent more slots than the target because of rejections and the empirical filter. Generation after the reset, FIRST (before training tasks).
NEW_FIRST_ID = 200
NEW_RUN_BASE = 440000 # 40 per slot, below 466560
NEW_KIND_SLOTS = [("INTF", "A", 10), ("TABL", "B", 11), ("STRU", "B", 3), ("MSAG", "D", 4), ("EXC", "D", 4)]
NEW_K = [("INTF", "free_text"), ("TABL", "incomplete"), ("TABL", "free_text"), ("MSAG", "free_text"), ("EXC", "incomplete")]
EXC_TOPICS = ["exception class (CX_...): a domain exception with context attributes and message texts",
"exception class (CX_...): an exception hierarchy with a common super class",
"exception class (CX_...): an exception that wraps a previous exception",
"exception class (CX_...): an exception with a message class and parameters in the text"]
def plan_new():
out, n = [], 0
for kind, cat, count in NEW_KIND_SLOTS:
for i in range(count):
out.append({"id": f"G{NEW_FIRST_ID + n:04d}", "kind": kind, "category": cat, "object_type": "CLAS" if kind == "EXC" else kind,
"topic": EXC_TOPICS[i % len(EXC_TOPICS)] if kind == "EXC" else None,
"difficulty": 3 if i % 3 == 2 else 2, "run_base": NEW_RUN_BASE + 40 * n})
n += 1
return out
def run_new(only, model, base_url):
"""Generate the new-kind eval candidates. The bundle is also checked against the training pool (no near duplicate of a training task)."""
train = overlap.load_pool("train")
for s in plan_new():
if only and s["id"] not in only:
continue
if os.path.exists(os.path.join(ROOT, "tasks_gen", "eval", s["id"], "generation.json")):
continue
avoid = [g for g in accepted_goals() if g][-170:]
topic = ((s["topic"] + ". ") if s["topic"] else "Choose a new, realistic business topic. ") + "Do not repeat these existing topics: " + "; ".join(avoid)
def extra(b, _t=train):
return [f"Too close to task {i} (spec {sc['spec']:.2f}, rules {sc['core']:.2f}, names {sc['name']:.2f}). Choose another topic and other object names."
for i, sc in overlap.check(overlap.load_bundle(b), _t)[:3]]
try:
log = generate(s["id"], "eval", s["object_type"], s["category"], s["difficulty"], model, base_url, s["run_base"], topic, extra_check=extra)
except BudgetExceeded as e:
print("BUDGET", e, flush=True)
break
except Exception as e: # noqa: BLE001
log = {"id": s["id"], "error": str(e)[:500]}
log.update(kind=s["kind"], spent_total=spent())
print(json.dumps(log), flush=True)
def run_new_k(model, base_url):
"""K variants (free text or incomplete spec, EPOD tool names) of the first accepted eval task of each new kind."""
plan = plan_new()
first_k = NEW_FIRST_ID + len(plan)
for j, (kind, style) in enumerate(NEW_K):
src = next((s["id"] for s in plan if s["kind"] == kind and os.path.exists(os.path.join(ROOT, "tasks_gen", "eval", s["id"], "generation.json"))
and json.load(open(os.path.join(ROOT, "tasks_gen", "eval", s["id"], "generation.json"))).get("accepted")), None)
new_id = f"G{first_k + j:04d}"
if not src or os.path.exists(os.path.join(ROOT, "tasks_gen", "eval", new_id, "generation.json")):
continue
try:
log = make_k_variant(src, new_id, style, model, base_url, NEW_RUN_BASE + 40 * (len(plan) + j), tool_schema=None, pool="eval")
except BudgetExceeded as e:
print("BUDGET", e, flush=True)
break
print(json.dumps(log), flush=True)
def accepted_goals(): def accepted_goals():
"""Goal lines of all generated tasks (eval and train), so the model does not repeat a topic.""" """Goal lines of all generated tasks (eval and train), so the model does not repeat a topic."""
out = [] out = []
@@ -90,6 +158,18 @@ def plan():
def main(): def main():
load_env(os.path.join(ROOT, ".env")) load_env(os.path.join(ROOT, ".env"))
cmd = sys.argv[1] if len(sys.argv) > 1 else "plan" cmd = sys.argv[1] if len(sys.argv) > 1 else "plan"
model, base_url = "deepseek-v4.1-flash:cloud", os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1")
if cmd == "plan-new":
for x in plan_new():
print(x)
print(len(plan_new()), "slots,", len(NEW_K), "K variants after them")
return
if cmd == "run-new":
run_new(set(sys.argv[2:]), model, base_url)
return
if cmd == "run-new-k":
run_new_k(model, base_url)
return
slots = plan() slots = plan()
if cmd == "plan": if cmd == "plan":
for s in slots: for s in slots:

200
harness/owntests.py Normal file
View File

@@ -0,0 +1,200 @@
"""Own-test mutation score of an accepted trajectory (Opus item D, 2026-10-06): metadata only, the acceptance filter does not change.
The model's own unit tests (a testclasses include of the contract class, or global test classes) are run against the faulty references
of the task (`faulty/`: mutants of the reference that the hidden tests kill). Per trajectory: one run with the correct reference (the tests must
pass there, otherwise they encode model specific behavior and a kill proves nothing), then one run per mutant. Status per mutant:
killed (an own test fails), survived (all own tests pass), invalid (the mutant or the tests do not activate, or no own test ran).
score = killed / (killed + survived) over the valid mutants. PROG tasks are not supported (the tests live inside the program source).
python3 -m harness.owntests [--tasks G1000 ...] [--limit N] [--workers 1] writes runs/traj/<run>/own_test_mutation.json
"""
import argparse
import glob
import json
import os
import re
import shutil
import sys
import time
from .adt_client import load_env
from .agents import OracleAgent
from . import mix
from .runner import Runner
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
POOL = os.path.join(ROOT, "tasks_gen", "train")
WORK = os.path.join(ROOT, "runs", "owntests")
RUN_BASE = 480000 # + sequence number; below 466560 is NOT needed here: the run numbers stay under 466560 with 4-char base36
RUN_BASE = 420000
MAX_MUTANTS = 5
def calls_of(messages):
res = {m.get("tool_call_id"): m["content"] for m in messages if m["role"] == "tool"}
out = []
for m in messages:
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
out.append((c["function"]["name"], a, res.get(c.get("id"), "")))
return out
def model_tests(rec, task_meta):
"""{'include': {OBJECT: source}, 'global': {NAME: source}} of the model's last successful pushes."""
contract = {c["name"].upper() for c in task_meta.get("contract", [])}
last = {}
for tool, a, res in calls_of(rec["messages"]):
if tool == "sap_push_source" and a.get("source") and '"success":true' in (res or "").replace(" ", ""):
last[(str(a.get("objectName", "")).upper(), str(a.get("includeType") or "").lower(), a.get("objectType"))] = a["source"]
pre = rec["prefix"].upper()
seed_hidden = {o["name"].replace("{{P}}", pre).upper() for k in ("seed", "hidden_tests") for o in task_meta.get(k, [])}
out = {"include": {}, "global": {}}
for (name, inc, otype), src in last.items():
contract_names = {c.replace("{{P}}", pre).upper() for c in contract}
if inc == "testclasses" and name in contract_names:
out["include"][name] = src
elif otype == "CLAS" and not inc and name not in contract_names and name not in seed_hidden \
and re.search(r"FOR\s+TESTING", src, re.I) and re.search(r"^\s*CLASS\s+\S+\s+DEFINITION[^.]*FOR\s+TESTING", src, re.I | re.M):
out["global"][name] = src
return out
def placeholder(src, prefix):
return re.sub(re.escape(prefix), "{{P}}", re.sub(re.escape(prefix.lower()), "{{p}}", src), flags=re.I) if False else \
src.replace(prefix.upper(), "{{P}}").replace(prefix.lower(), "{{p}}")
def mutant_files(task_id):
"""[(reference object name, path)] of faulty/ (a K variant has none: the base task's)."""
meta = json.load(open(os.path.join(POOL, task_id, "task.json")))
d = os.path.join(POOL, meta.get("base_task") or task_id, "faulty")
return sorted(glob.glob(os.path.join(d, "m*_*")))
def derive(task_id, run_label, tests, prefix, mutant_path=None):
"""Build a temporary task: the reference (optionally with one mutated object) plus the model's own tests only."""
src_dir = os.path.join(POOL, task_id)
pool = os.path.join(WORK, "pool")
dst = os.path.join(pool, task_id)
shutil.rmtree(dst, ignore_errors=True)
shutil.copytree(src_dir, dst, ignore=shutil.ignore_patterns("faulty", "generation.json", "review.json", "mutation.json", "empirical.json"))
meta = json.load(open(os.path.join(dst, "task.json")))
pre = prefix.upper()
mutated = None
if mutant_path:
base = re.sub(r"^m\d+_", "", os.path.basename(mutant_path))
refs = []
for o in meta["reference"]:
f = o.get("file")
is_test = o["type"] == "CLAS" and f and re.search(r"^\s*CLASS\s+\S+\s+DEFINITION[^.]*FOR\s+TESTING",
open(os.path.join(dst, f)).read(), re.I | re.M)
if is_test:
continue # the reference's own global test class: not the model's
o = dict(o)
o.pop("testclasses_file", None) # the reference's local tests are out; the model's go in
if mutant_path and f and os.path.basename(f) == base:
shutil.copy(mutant_path, os.path.join(dst, f))
mutated = o["name"]
name_up = o["name"].replace("{{P}}", pre).upper()
if name_up in tests["include"]:
rel = "reference/_own_include_%s.abap" % re.sub(r"\W", "_", name_up)
open(os.path.join(dst, rel), "w").write(placeholder(tests["include"][name_up], prefix))
o["testclasses_file"] = rel
refs.append(o)
for i, (name, src) in enumerate(tests["global"].items()):
rel = "reference/_own_global_%d.clas.abap" % i
open(os.path.join(dst, rel), "w").write(placeholder(src, prefix))
refs.append({"type": "CLAS", "name": placeholder(name, prefix), "file": rel, "description": "own test class"})
meta["reference"] = refs
meta["budget"] = dict(meta.get("budget", {}), max_tool_calls=200)
json.dump(meta, open(os.path.join(dst, "task.json"), "w"), indent=1)
return pool, mutated
def run_one(pool, task_id, run_no):
runner = Runner(pool, os.path.join(WORK, "runs"))
rep, _ = runner.run(task_id, OracleAgent(), run_no, teardown=True)
own = rep.get("own_tests") or {}
g = rep.get("gates") or {}
return {"tests": own.get("tests", 0), "failures": own.get("failures", 0), "active": bool(g.get("G1_active")), "run": run_no}
def score_run(run_dir, seq):
rec = json.load(open(os.path.join(run_dir, "record.json")))
tid = rec["task"]["id"]
meta = json.load(open(os.path.join(POOL, tid, "task.json")))
out = {"task": tid, "run": rec["run"], "kind": mix.kind_of_task_dir(tid), "time": time.strftime("%F %T")}
if any(c.get("type") == "PROG" for c in meta.get("contract", [])):
return dict(out, status="not_supported", reason="PROG: the tests are inside the program source")
tests = model_tests(rec, meta)
if not tests["include"] and not tests["global"]:
return dict(out, status="no_own_tests", score=None, mutants=[])
muts = mutant_files(tid)[:MAX_MUTANTS]
if not muts:
return dict(out, status="no_mutants", score=None, mutants=[])
pre = rec["prefix"]
pool, _ = derive(tid, "base", tests, pre)
base = run_one(pool, tid, RUN_BASE + seq * 10)
out["reference_run"] = base
out["tests_pass_on_reference"] = base["active"] and base["tests"] > 0 and base["failures"] == 0
res = []
for k, mp in enumerate(muts):
pool, mutated = derive(tid, "m%d" % k, tests, pre, mp)
r = run_one(pool, tid, RUN_BASE + seq * 10 + 1 + k)
status = "invalid" if (not r["active"] or r["tests"] == 0) else ("killed" if r["failures"] > 0 else "survived")
res.append({"mutant": os.path.basename(mp), "object": mutated, "status": status, "tests": r["tests"], "failures": r["failures"]})
valid = [x for x in res if x["status"] != "invalid"]
killed = [x for x in res if x["status"] == "killed"]
out.update(status="scored", mutants=res, valid=len(valid), killed=len(killed),
score=round(len(killed) / len(valid), 2) if valid else None)
shutil.rmtree(os.path.join(WORK, "pool", tid), ignore_errors=True)
return out
def main():
load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser()
ap.add_argument("--tasks", nargs="*")
ap.add_argument("--limit", type=int)
ap.add_argument("--redo", action="store_true")
a = ap.parse_args()
os.makedirs(WORK, exist_ok=True)
rows = [json.loads(l) for l in open(os.path.join(ROOT, "runs", "traj", "summary.jsonl"))]
todo = []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-")
if not os.path.exists(os.path.join(p, "record.json")):
continue
if a.tasks and r["task"] not in a.tasks:
continue
if os.path.exists(os.path.join(p, "own_test_mutation.json")) and not a.redo:
continue
rec = json.load(open(os.path.join(p, "record.json")))
if acc.judge(rec, r, 80)[0]:
todo.append((r, p))
if a.limit:
todo = todo[:a.limit]
print(len(todo), "accepted trajectories to score", flush=True)
for i, (r, p) in enumerate(todo):
t0 = time.time()
try:
res = score_run(p, i)
except Exception as e: # noqa: BLE001
res = {"task": r["task"], "status": "error", "error": repr(e)[:300]}
json.dump(res, open(os.path.join(p, "own_test_mutation.json"), "w"), indent=1)
print(r["task"], res.get("status"), res.get("score"), "valid", res.get("valid"), "killed", res.get("killed"),
"ref_ok", res.get("tests_pass_on_reference"), "%.0fs" % (time.time() - t0), flush=True)
if __name__ == "__main__":
main()

View File

@@ -51,6 +51,9 @@ def activation_messages(text, is_error=False):
RUN_PREFIX = re.compile(r"^Z\d[0-9A-Z]{5,6}_", re.I) RUN_PREFIX = re.compile(r"^Z\d[0-9A-Z]{5,6}_", re.I)
# The model sometimes names its own helper or test class with the prefix inside the name (ZCL_Z4CGT1HJ_JOB_COST_TEST). Such objects of
# other runs stay behind in A4H and show up in lists and searches (57 of 90 accepted trajectories saw them, 2026-10-06).
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
class BudgetExceeded(Exception): class BudgetExceeded(Exception):
@@ -111,7 +114,10 @@ class ToolProxy:
def _foreign(self, name): def _foreign(self, name):
n = (name or "").upper() n = (name or "").upper()
return bool(RUN_PREFIX.match(n)) and not n.startswith(self.prefix) if RUN_PREFIX.match(n):
return not n.startswith(self.prefix)
m = MID_PREFIX.match(n)
return bool(m) and m.group(1).upper() != self.prefix
def _filter(self, tool, text): def _filter(self, tool, text):
if tool == "sap_short_dumps": # only dumps of this run: other runs (and mutants) also write dumps if tool == "sap_short_dumps": # only dumps of this run: other runs (and mutants) also write dumps
@@ -124,7 +130,7 @@ class ToolProxy:
data["count"] = len(data["dumps"]) data["count"] = len(data["dumps"])
return json.dumps(data) return json.dumps(data)
return text return text
if tool not in ("sap_search_object", "sap_usage_references"): if tool not in ("sap_search_object", "sap_usage_references", "sap_inactive_objects"):
return text return text
try: try:
data = json.loads(text) data = json.loads(text)
@@ -148,7 +154,9 @@ class ToolProxy:
else: else:
self.calls += 1 self.calls += 1
name = str(args.get("objectName", "")).upper() name = str(args.get("objectName", "")).upper()
if tool in WRITE_TOOLS and self._foreign(name): if (tool in WRITE_TOOLS or tool in ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info",
"sap_run_unit_test", "sap_check_object", "sap_syntax_check", "sap_atc_run")) \
and self._foreign(name):
result = (True, f"{name} is not available.") result = (True, f"{name} is not available.")
else: else:
if tool in WRITE_TOOLS and (tool == "sap_activate" or args.get("activate", True)): if tool in WRITE_TOOLS and (tool == "sap_activate" or args.get("activate", True)):

View File

@@ -153,9 +153,16 @@ class Runner:
return out return out
def _objects_with_prefix(self, mcp, prefix): def _objects_with_prefix(self, mcp, prefix):
_, text = mcp.call("sap_search_object", {"query": prefix + "*", "maxResults": 200}) """Objects of the run: the name starts with the prefix, or contains it after a short type part (the model named its own
return [d for d in (_json(text) or []) if d.get("name", "").upper().startswith(prefix) helper or test class ZCL_<prefix>_..., 27 such objects were left behind before 2026-10-06)."""
and d.get("objectType")] # skips STOB entries of CDS entities out = {}
for q in (prefix + "*", "*_" + prefix + "*"):
_, text = mcp.call("sap_search_object", {"query": q, "maxResults": 200})
for d in (_json(text) or []):
n = d.get("name", "").upper()
if d.get("objectType") and (n.startswith(prefix) or re.match(r"^[A-Z]{1,5}_" + re.escape(prefix), n)):
out[n] = d # skips STOB entries of CDS entities (no objectType)
return list(out.values())
def _source(self, mcp, otype, name, fg=None): def _source(self, mcp, otype, name, fg=None):
err, text = mcp.call("sap_pull_source", _obj_args(otype, name, fg)) err, text = mcp.call("sap_pull_source", _obj_args(otype, name, fg))

89
harness/sweep.py Normal file
View File

@@ -0,0 +1,89 @@
"""Find (and delete) objects that harness runs left in A4H: names with a run prefix at the start (Z4CGT1HJ_X) or inside (ZCL_Z4CGT1HJ_X).
python3 -m harness.sweep dry run: list them by prefix
python3 -m harness.sweep --delete delete them (ADT deletion API, dependency order); a locked object is reported, not forced
python3 -m harness.sweep --keep-recent 60 do not touch prefixes of runs whose directory changed in the last 60 minutes (default 60)
Never run it while a controller or a series is writing: it can only judge by directory age.
"""
import argparse
import glob
import json
import os
import re
import time
from .adt_client import load_env
from .mcp_client import McpClient
from .runner import DELETE_ORDER, delete_uris
from .task import prefix_for
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
START = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
MID = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
def recent_prefixes(minutes):
keep = set()
for d in glob.glob(os.path.join(ROOT, "runs", "**", "*_G*_*"), recursive=True) + glob.glob(os.path.join(ROOT, "runs", "**", "*_T*_*"), recursive=True):
b = os.path.basename(d)
m = re.match(r"^(\d+)_([GT]\d+)_", b)
if m and time.time() - os.path.getmtime(d) < minutes * 60:
keep.add(prefix_for(int(m.group(1)), m.group(2)).upper())
return keep
def main():
load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser()
ap.add_argument("--delete", action="store_true")
ap.add_argument("--keep-recent", type=int, default=60)
a = ap.parse_args()
found = {}
with McpClient() as m:
# run prefixes are Z + a digit + 6 characters: one query per digit (a single "Z*" query is cut at 1000 results), at the start and after the type part
for q in [f"Z{d}*" for d in "0123456789"] + [f"*_Z{d}*" for d in "0123456789"] + ["ZPROBE*", "ZTEST*"]:
try:
for o in json.loads(m.call("sap_search_object", {"query": q, "maxResults": 1000})[1]):
if o.get("objectType"): # a STOB entry of a CDS view has the same name and no objectType: skip it
found[(o["name"], o["objectType"])] = o
except ValueError:
pass
keep = recent_prefixes(a.keep_recent)
by_prefix = {}
for (n, _t), o in found.items():
mm = START.match(n.upper()) or MID.match(n.upper())
if not mm or not o.get("objectType"):
continue
pre = mm.group(1).upper()
if len(pre) != 9 or not pre[1].isdigit():
continue
if pre in keep:
continue
by_prefix.setdefault(pre, []).append(o)
total = sum(len(v) for v in by_prefix.values())
print(f"{total} objects of {len(by_prefix)} run prefixes ({len(keep)} recent prefixes kept)")
for pre, objs in sorted(by_prefix.items()):
print(" ", pre, [o["name"] for o in objs][:6])
json.dump({p: [o["name"] for o in v] for p, v in by_prefix.items()}, open(os.path.join(ROOT, "runs", "sweep.json"), "w"), indent=1)
if not a.delete:
return
objs = sorted([o for v in by_prefix.values() for o in v],
key=lambda o: DELETE_ORDER.index(o["objectType"]) if o["objectType"] in DELETE_ORDER else 99)
done = failed = 0
for i in range(0, len(objs), 10):
uris = []
for o in objs[i:i + 10]:
uris.append(o["uri"])
res = delete_uris(uris)
for u, v in res.items():
if v["deleted"]:
done += 1
else:
failed += 1
print("not deleted:", u.split("/")[-1], v["msg"])
print(f"deleted {done}, not deleted {failed}")
if __name__ == "__main__":
main()

View File

@@ -28,7 +28,20 @@ POOL = os.path.join(ROOT, "tasks_gen", "train")
OUT = os.path.join(ROOT, "runs", "traj") OUT = os.path.join(ROOT, "runs", "traj")
RUN_BASE = 200000 # 200000 + (task number - 1000) * 3 + attempt (a digit must lead the 4-char base36 run: < 466560) RUN_BASE = 200000 # 200000 + (task number - 1000) * 3 + attempt (a digit must lead the 4-char base36 run: < 466560)
MODEL = "deepseek-v4.1-flash:cloud" MODEL = "deepseek-v4.1-flash:cloud"
CDS_CALLS = 100 # tool-call budget for tasks with a CDS contract object (eval keeps 60; Kral 2026-10-05) # empty_response fix (2026-10-06, docs/empty-response.md): 13 of 119 runs ended with one turn that used the whole output limit on
# reasoning. Cap per turn 24000 (only 3 of 1997 turns were legitimately longer), one retry at another temperature (the same
# sample ran away again in 7 of 13 runs), and the stream guard (a streamed turn with only reasoning is cut after that many reasoning
# tokens) which stays off until it is verified on the cloud model (env STREAM_GUARD, for example 9000).
TEACHER_MAX_TOKENS = 24000
EMPTY_RETRIES = 1
RETRY_TEMPERATURE = 0.8
STREAM_GUARD = int(os.environ["STREAM_GUARD"]) if os.environ.get("STREAM_GUARD") else None
CDS_CALLS = 100
def new_agent():
return LlmAgent(MODEL, loop_guard=3, max_tokens=TEACHER_MAX_TOKENS, empty_retries=EMPTY_RETRIES,
retry_temperature=RETRY_TEMPERATURE, stream_guard=STREAM_GUARD) # tool-call budget for tasks with a CDS contract object (eval keeps 60; Kral 2026-10-05)
LOCK = threading.Lock() LOCK = threading.Lock()
@@ -74,7 +87,7 @@ def one(task_id, attempt, stop):
print("BUDGET", e, flush=True) print("BUDGET", e, flush=True)
return return
run_no = RUN_BASE + (int("".join(c for c in task_id if c.isdigit())) - 1000) * 3 + attempt run_no = RUN_BASE + (int("".join(c for c in task_id if c.isdigit())) - 1000) * 3 + attempt
agent = LlmAgent(MODEL, loop_guard=3) agent = new_agent()
runner = Runner(POOL, OUT) runner = Runner(POOL, OUT)
t0 = time.time() t0 = time.time()
row = {"task": task_id, "attempt": attempt, "run": run_no, "model": MODEL} row = {"task": task_id, "attempt": attempt, "run": run_no, "model": MODEL}
@@ -97,7 +110,7 @@ def one(task_id, attempt, stop):
for d in glob.glob(os.path.join(OUT, f"{run_no}_{task_id}_*")): for d in glob.glob(os.path.join(OUT, f"{run_no}_{task_id}_*")):
os.makedirs(os.path.join(OUT, "_aborted"), exist_ok=True) os.makedirs(os.path.join(OUT, "_aborted"), exist_ok=True)
os.rename(d, os.path.join(OUT, "_aborted", os.path.basename(d) + "_" + str(int(time.time())))) os.rename(d, os.path.join(OUT, "_aborted", os.path.basename(d) + "_" + str(int(time.time()))))
agent = LlmAgent(MODEL, loop_guard=3) agent = new_agent()
score = (rep.get("score") or {}).get("total") score = (rep.get("score") or {}).get("total")
rec = os.path.exists(os.path.join(run_dir, "record.json")) rec = os.path.exists(os.path.join(run_dir, "record.json"))
row.update(score=score, setup_failed=bool(rep.get("setup_failed")), end_reason=rep.get("end_reason"), row.update(score=score, setup_failed=bool(rep.get("setup_failed")), end_reason=rep.get("end_reason"),

22
scripts_probe/after_owntests.sh Executable file
View File

@@ -0,0 +1,22 @@
#!/bin/sh
# Waits for harness.owntests to end, then: report, rebuild the stage 2 set with the own-test score as metadata, update the HF dataset, commit.
cd "$HOME/projects/abap-llm/harness" || exit 1
PID="$1"
while kill -0 "$PID" 2>/dev/null; do sleep 60; done
python3 train/own_test_report.py > runs/own_test_report.log 2>&1
train/.venv/bin/python train/build_stage2.py --out runs/stage2_data --hook hooks_example:own_test_weight > runs/build_after_owntests.log 2>&1
python3 train/build_doc.py >> runs/build_after_owntests.log 2>&1
set -a; . ./.env; set +a
train/.venv/bin/python - >> runs/build_after_owntests.log 2>&1 <<'PY'
import os
from huggingface_hub import HfApi
api = HfApi(token=os.environ["HF_TOKEN"])
for f in ("stage2_train.jsonl", "stage2_valid.jsonl", "stage2_reserve.jsonl", "build_report.json", "README.md"):
api.upload_file(path_or_fileobj="runs/stage2_data/" + f, path_in_repo=f, repo_id="erhankeseli/abap-stage2-data", repo_type="dataset")
print("uploaded")
PY
export GIT_AUTHOR_NAME=Kral GIT_AUTHOR_EMAIL=kral@local GIT_COMMITTER_NAME=Kral GIT_COMMITTER_EMAIL=kral@local
git add docs train && git commit -q -m "D: own-test mutation scores of the accepted trajectories (metadata only), stage 2 set rebuilt with the score
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>"
echo "after_owntests done $(date)" >> runs/own_test_report.log

View File

@@ -21,7 +21,7 @@ CLASS zprobe0eq_locks IMPLEMENTATION.
TABLES enq = lt_enq TABLES enq = lt_enq
EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3. EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.
DATA(lv_n) = 0. DATA(lv_n) = 0.
LOOP AT lt_enq ASSIGNING FIELD-SYMBOL(<l>) WHERE garg CS 'ZPROBE0' OR garg CS 'Z4AJ' OR garg CS 'Z4AE'. LOOP AT lt_enq ASSIGNING FIELD-SYMBOL(<l>).
lv_n = lv_n + 1. lv_n = lv_n + 1.
DATA(lv_line) = ||. DATA(lv_line) = ||.
DO. DO.
@@ -35,15 +35,16 @@ CLASS zprobe0eq_locks IMPLEMENTATION.
ENDDO. ENDDO.
out->write( lv_line ). out->write( lv_line ).
ENDLOOP. ENDLOOP.
out->write( |enqueue entries matching the probe prefixes: { lv_n } of { lines( lt_enq ) } (subrc { lv_subrc })| ). out->write( |enqueue entries listed: { lv_n } of { lines( lt_enq ) } (subrc { lv_subrc })| ).
ENDMETHOD. ENDMETHOD.
ENDCLASS. ENDCLASS.
""" """
def ensure_reader(m): def ensure_reader(m, force=False):
e, t = m.call("sap_search_object", {"query": READER}) e, t = m.call("sap_search_object", {"query": READER})
if READER not in t: if READER not in t or force:
m.call("sap_create_object", {"objectType": "CLAS", "objectName": READER, "packageName": "$TMP", "description": "enqueue reader probe"}) if READER not in t:
m.call("sap_create_object", {"objectType": "CLAS", "objectName": READER, "packageName": "$TMP", "description": "enqueue reader probe"})
e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": READER, "source": READER_SRC}) e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": READER, "source": READER_SRC})
if '"success":true' not in t.replace(" ", ""): if '"success":true' not in t.replace(" ", ""):
print("reader activation:", t[:600]) print("reader activation:", t[:600])
@@ -54,5 +55,5 @@ def read_locks(m):
if __name__ == "__main__": if __name__ == "__main__":
with McpClient() as m: with McpClient() as m:
ensure_reader(m) ensure_reader(m, force=True)
print(read_locks(m)[:2500]) print(read_locks(m)[:3500])

View File

@@ -180,3 +180,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- 2026-10-05 22:35 **pipeline stopped at 21:43 by an infrastructure outage, found 22:26 (my miss).** The MCP server (127.0.0.1:3000) refused connections for a short time; three runs got `URLError(ConnectionRefusedError)` and the pipeline counted three equal harness errors and stopped (rule). Not a model or harness bug: A4H was up (11 h), the MCP server answers again. New `harness/infra.py`: an outage of MCP or A4H is waited out (check every 30 s, two answers in a row), the run's objects are deleted, the same run starts again; it is not counted as an event or a failure (`runs/traj/outages.jsonl`). Same for series A (applies after a restart of `harness.localqwen`; the running series keeps the old code and would skip the task). The 3 interrupted runs were cleaned (11 objects) and get their second attempt. Budget: panel 50.84 at ledger 113.36 (2.54 ledger per usage in the last stretch); guard recomputed: panel left 9.16 - 3.0 reserve = 6.2 usage x 1.2 = 7.4 ledger, `BUDGET_LIMIT_USD` 129 (guard 121). Dashboard: a stopped pipeline shows no burn rate or ETA. - 2026-10-05 22:35 **pipeline stopped at 21:43 by an infrastructure outage, found 22:26 (my miss).** The MCP server (127.0.0.1:3000) refused connections for a short time; three runs got `URLError(ConnectionRefusedError)` and the pipeline counted three equal harness errors and stopped (rule). Not a model or harness bug: A4H was up (11 h), the MCP server answers again. New `harness/infra.py`: an outage of MCP or A4H is waited out (check every 30 s, two answers in a row), the run's objects are deleted, the same run starts again; it is not counted as an event or a failure (`runs/traj/outages.jsonl`). Same for series A (applies after a restart of `harness.localqwen`; the running series keeps the old code and would skip the task). The 3 interrupted runs were cleaned (11 objects) and get their second attempt. Budget: panel 50.84 at ledger 113.36 (2.54 ledger per usage in the last stretch); guard recomputed: panel left 9.16 - 3.0 reserve = 6.2 usage x 1.2 = 7.4 ledger, `BUDGET_LIMIT_USD` 129 (guard 121). Dashboard: a stopped pipeline shows no burn rate or ETA.
- 2026-10-05 22:40 cause of the 21:43 outage confirmed by Kral: he restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment; their objects could be deleted afterwards (no stale lock), so a restart of the MCP server in the middle of a write did not leak a lock in this case. Rule: restarting Eclipse is fine now (the pipeline waits and reruns), but check `python3 scripts_probe/lockprobe.py` for stale locks afterwards. - 2026-10-05 22:40 cause of the 21:43 outage confirmed by Kral: he restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment; their objects could be deleted afterwards (no stale lock), so a restart of the MCP server in the middle of a write did not leak a lock in this case. Rule: restarting Eclipse is fine now (the pipeline waits and reruns), but check `python3 scripts_probe/lockprobe.py` for stale locks afterwards.
- 2026-10-06 06:40 **correction of the 22:40 entry and a new finding.** (1) An Eclipse restart in the middle of a write does leak a lock: `ZCL_Z4CGT1HJ_JOB_COST_TEST` (enqueue SEOCLSENQ, mode X, created 21:46, three minutes after the 21:43 restart) stayed locked until the next morning; details in `docs/epod-lock-leak.md` (case 3). I had missed it because the teardown only cleaned names that start with the run prefix. SM12 is needed (Kral). (2) **Leftover objects with the prefix inside the name** (ZCL_<prefix>_..., the model's own test classes): 27 stayed in A4H, and **57 of 90 accepted trajectories contain other runs' object names in tool results** (mostly `sap_inactive_objects`, in 30 cases the model pulled another run's class). Fixed: `harness/proxy.py` hides them (search, usage, inactive list) and answers "not available" for reads and writes of another run's object; `harness/runner.py` finds mid-name objects for teardown; `harness/sweep.py` deleted the leftovers (28 objects; the locked class remains). The 90 trajectories are not changed yet: see the data scrub option of `train/build_stage2.py`.

236
train/analyze.py Normal file
View File

@@ -0,0 +1,236 @@
"""Analysis of the accepted DeepSeek trajectories (Opus item B, 2026-10-06): repair taxonomy, near duplicates,
empty_response, and the behaviors the teacher shows and Qwen lacks.
python3 train/analyze.py writes runs/analysis/analysis.json and prints the numbers
"""
import glob
import json
import os
import re
import statistics
import sys
from collections import Counter
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
from harness import mix # noqa: E402
from harness.generator import RESERVED # noqa: E402
WRITE = ("sap_push_source", "sap_push_element", "sap_push_message", "sap_activate", "sap_create_object")
SEARCH_LIKE = ("sap_search_object", "sap_pull_source", "sap_object_members", "sap_object_structure", "sap_element_info",
"sap_usage_references", "sap_sql_query", "sap_inactive_objects")
SIG = re.compile(r"(?:IMPORTING|EXPORTING|RETURNING|CHANGING)[^.]*?\bTYPE\s+(?:c|n|p|x)\s+LENGTH\b|\bVALUE\([^)]*\)\s+TYPE\s+(?:c|n|p|x)\s+LENGTH\b", re.I | re.S)
def pct(v, p):
v = sorted(v)
return v[min(len(v) - 1, int(p * len(v)))] if v else None
def calls_of(messages):
"""[(index, tool name, args dict, result text)] in order."""
res = {m.get("tool_call_id"): m["content"] for m in messages if m["role"] == "tool"}
out = []
for i, m in enumerate(messages):
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
out.append((i, c["function"]["name"], a, res.get(c.get("id"), "")))
return out
def failed(tool, text):
t = (text or "").replace(" ", "")
return tool in WRITE and (t.startswith("ERROR:") or '"success":false' in t)
def error_text(text):
try:
o = json.loads(text.replace("ERROR: ", "", 1) if text.startswith("ERROR:") else text)
except ValueError:
return text[:200]
if not isinstance(o, dict):
return text[:200]
msgs = [m.get("message", "") for m in (o.get("activation") or {}).get("messages", []) if m.get("severity") == "E"]
out = " | ".join(msgs) or str(o.get("error") or o.get("message") or "")
sc = o.get("syntaxCheck")
if sc:
out += " | syntaxCheck: " + "; ".join(m.get("message", "") for m in sc.get("messages", []))
return out[:400]
def classify(src, err):
"""Labels of a failed write (a write can match more than one)."""
labels = []
if src and SIG.search(src):
labels.append("type_c_length_in_signature")
low = (err or "").lower()
if "longer than the allowed 30 characters" in low or "30 characters" in low:
labels.append("name_over_30")
words = {w for w in re.findall(r"[A-Za-z_]+", err or "")}
if "is not valid" in low or "reserved" in low or any(w.upper() in RESERVED and w.upper() in (err or "").upper() for w in words if len(w) > 3 and w.isupper()):
if re.search(r"\b(?:" + "|".join(sorted(RESERVED, key=len, reverse=True)) + r")\b", err or "", re.I) or "reserved" in low:
labels.append("reserved_word")
return labels
def norm(err):
e = re.sub(r"\bZ[0-9A-Z]{7}_\w+", "<OBJ>", err or "")
e = re.sub(r"\b[A-Z][A-Z0-9_]{5,}\b", "<NAME>", e)
e = re.sub(r"\d+", "N", e)
return e[:110]
def behaviors(messages):
cs = calls_of(messages)
out = {"first_write": None, "searches_before_first_write": 0, "max_consecutive_search_object": 0, "calls": len(cs),
"failed_writes": 0, "repaired": False, "calls_to_repair": None, "same_push_repeats": 0}
run_so = 0
first_fail = None
last_src = {}
for n, (i, tool, a, res) in enumerate(cs):
if tool == "sap_search_object":
run_so += 1
out["max_consecutive_search_object"] = max(out["max_consecutive_search_object"], run_so)
else:
run_so = 0
if tool in ("sap_push_source", "sap_push_element", "sap_push_message") and out["first_write"] is None:
out["first_write"] = n
if out["first_write"] is None and tool in SEARCH_LIKE:
out["searches_before_first_write"] += 1
if tool == "sap_push_source":
key = (a.get("objectName"), a.get("includeType"))
if last_src.get(key) == a.get("source") and a.get("source"):
out["same_push_repeats"] += 1
last_src[key] = a.get("source")
if failed(tool, res):
out["failed_writes"] += 1
if first_fail is None:
first_fail = (n, a.get("objectName"), a.get("source"))
elif first_fail and tool in ("sap_push_source", "sap_push_element") and '"success":true' in (res or "").replace(" ", "") \
and a.get("objectName") == first_fail[1] and a.get("source") != first_fail[2] and not out["repaired"]:
out["repaired"] = True
out["calls_to_repair"] = n - first_fail[0]
return out
def shingles(text, k=5):
t = re.sub(r"\bz[0-9a-z]{7}_", "", text.lower())
w = re.findall(r"\w+", t)
return {" ".join(w[i:i + k]) for i in range(max(len(w) - k + 1, 0))}
def final_sources(messages):
"""Last pushed source per (object, include): the trajectory's own code."""
last = {}
for i, tool, a, res in calls_of(messages):
if tool == "sap_push_source" and a.get("source") and '"success":true' in (res or "").replace(" ", ""):
last[(a.get("objectName"), a.get("includeType"))] = a["source"]
return "\n".join(last.values())
def main():
rows = [json.loads(l) for l in open(os.path.join(ROOT, "runs", "traj", "summary.jsonl"))]
acc_recs, empty_runs, all_turn_out = [], [], []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
if not os.path.exists(p):
continue
rec = json.load(open(p))
ok, why = acc.judge(rec, r, 80)
md = rec["metadata"]
if ok:
acc_recs.append((r, rec))
all_turn_out += [u.get("completion_tokens", 0) for u in md.get("turn_usage", [])]
if md.get("end_reason") == "empty_response":
empty_runs.append((r, rec))
out = {"accepted": len(acc_recs), "runs": len(rows)}
# ---- repair taxonomy
fail_ct, labels_ct, traj_with, repaired_with, top = 0, Counter(), Counter(), Counter(), Counter()
per_label_examples = {}
kinds = Counter()
for r, rec in acc_recs:
seen = {}
for i, tool, a, res in calls_of(rec["messages"]):
if failed(tool, res):
fail_ct += 1
err = error_text(res)
top[norm(err)] += 1
for l in classify(a.get("source"), err):
labels_ct[l] += 1
seen[l] = True
per_label_examples.setdefault(l, (r["task"], err[:160]))
b = behaviors(rec["messages"])
for l in seen:
traj_with[l] += 1
if b["repaired"]:
repaired_with[l] += 1
out["repair_taxonomy"] = {"failed_writes_in_accepted": fail_ct, "step3_errors": {
l: {"failed_writes": labels_ct[l], "trajectories": traj_with[l], "trajectories_with_repair": repaired_with[l],
"example": per_label_examples.get(l)} for l in ("type_c_length_in_signature", "reserved_word", "name_over_30")},
"top_error_messages": top.most_common(15)}
# ---- teacher behaviors vs Qwen
def stats(recs):
b = [behaviors(x["messages"]) for x in recs]
had_fail = [x for x in b if x["failed_writes"]]
sb = [x["searches_before_first_write"] for x in b if x["first_write"] is not None]
return {"n": len(b), "with_a_failed_write": len(had_fail),
"repair_after_first_error": sum(x["repaired"] for x in had_fail), "median_calls_to_repair":
statistics.median([x["calls_to_repair"] for x in had_fail if x["calls_to_repair"] is not None] or [None]),
"median_searches_before_first_write": statistics.median(sb) if sb else None, "p90_searches_before_first_write": pct(sb, .9),
"max_consecutive_search_object": max([x["max_consecutive_search_object"] for x in b] or [0]),
"runs_with_no_write": sum(1 for x in b if x["first_write"] is None),
"runs_with_same_push_repeat": sum(1 for x in b if x["same_push_repeats"] > 0)}
out["teacher"] = stats([rec for _, rec in acc_recs])
qrecs = []
for p in glob.glob(os.path.join(ROOT, "runs", "local_qwen", "runs", "*", "record.json")):
qrecs.append(json.load(open(p)))
out["qwen_series_A"] = stats(qrecs)
out["qwen_series_A"]["note"] = "all 20 runs of series A (accepted and failed)"
# ---- near duplicates
sh = [(r["task"], r["run"], shingles(final_sources(rec["messages"])), shingles(rec["messages"][-1].get("content") or "")) for r, rec in acc_recs]
dup = []
for i in range(len(sh)):
for j in range(i + 1, len(sh)):
a, b = sh[i], sh[j]
for k, nm in ((2, "code"), (3, "report")):
if a[k] and b[k]:
jac = len(a[k] & b[k]) / len(a[k] | b[k])
if jac >= (0.8 if nm == "code" else 0.85):
dup.append({"a": f"{a[0]}_r{a[1]}", "b": f"{b[0]}_r{b[1]}", "what": nm, "jaccard": round(jac, 2), "same_task": a[0] == b[0]})
out["near_duplicates"] = {"pairs_flagged": len(dup), "same_task_pairs": sum(1 for d in dup if d["same_task"]), "examples": dup[:12]}
# ---- empty_response
er = []
for r, rec in empty_runs:
md = rec["metadata"]
tu = md.get("turn_usage", [])
last = [u.get("completion_tokens", 0) for u in tu[-3:]]
cs = calls_of(rec["messages"])
lastcall = cs[-1] if cs else None
er.append({"task": r["task"], "kind": mix.kind_of_task_dir(r["task"]), "category": rec["task"]["category"], "calls": md.get("tool_calls"),
"turns": len(tu), "last3_output_tokens": last, "last_tool": lastcall[1] if lastcall else None,
"last_tool_result_chars": len(lastcall[3]) if lastcall else None,
"context_tokens_last": (tu[-1].get("prompt_tokens") if tu else None)})
out["empty_response"] = {"runs": len(er), "of": len(rows), "by_kind": Counter(e["kind"] for e in er),
"by_last_tool": Counter(e["last_tool"] for e in er), "detail": er,
"output_tokens_of_normal_turns": {"p50": pct(all_turn_out, .5), "p90": pct(all_turn_out, .9),
"p99": pct(all_turn_out, .99), "p999": pct(all_turn_out, .999), "max": max(all_turn_out or [0]),
"turns": len(all_turn_out), "turns_over_16k": sum(1 for x in all_turn_out if x > 16000),
"turns_over_20k": sum(1 for x in all_turn_out if x > 20000)}}
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "analysis.json"), "w"), indent=1, default=dict)
print(json.dumps(out, indent=1, default=dict)[:6500])
if __name__ == "__main__":
main()

91
train/build_doc.py Normal file
View File

@@ -0,0 +1,91 @@
"""Writes docs/stage2-build.md from runs/stage2_data/build_report.json (run after train/build_stage2.py)."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
rep = json.load(open(os.path.join(ROOT, "runs", "stage2_data", "build_report.json")))
tr, va, rs = rep["train"], rep["valid"], rep["reserve"]
s1 = rep["stage1_stats"]["train"]["tokens"]
s1d = rep["stage1_stats"]["train"]["docs"]
L2 = tr["loss_tokens"]
def row(e1, e2, L=L2):
a, b = s1 * e1, L * e2
return f"| {e1} | {e2} | {a / 1e6:.2f} M | {b / 1e6:.2f} M | {100 * b / (a + b):.0f} % | {(a + b) / 1e6:.1f} M |"
drop = rep["dropped"]
dropped_lines = "\n".join(f"- {k}: {v}" for k, v in drop.items())
kinds = "\n".join(f"| {k} | {v['samples']} | {v['families']} | {v['short']} |" for k, v in rep["kinds"].items())
doc = f"""# Stage 2 training set, build of {rep['built']}
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
## Result
{rep['steps']['accepted_trajectories']} accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> {rep['steps']['after_eval_overlap']} after the eval overlap check
-> {rep['steps']['after_token_limit']} after the 48k limit (none cut) -> CLAS cap 35 %: {rep['steps']['clas_cap']['clas_kept']} CLAS kept, **{rep['steps']['clas_cap']['clas_reserve']} CLAS in reserve** (`stage2_reserve.jsonl`)
-> **{tr['samples']} train + {va['samples']} valid samples** ({tr['tokens'] / 1e6:.2f} M + {va['tokens'] / 1e6:.2f} M tokens, {tr['loss_tokens'] / 1e6:.2f} M loss tokens in train, p50 {tr['p50']}, p95 {tr['p95']}, max {tr['max']}, repair share {tr['repair_share']}).
Dropped:
{dropped_lines}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
## Not good enough yet (kinds under the minimum of {rep['settings']['min_per_kind']})
| kind | samples | families | short |
|---|---|---|---|
{kinds}
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
{row(2, 3)}
{row(1, 3)}
{row(0.5, 3)}
{row(0.33, 3)}
{row(1, 3, 3 * L2)}
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
| class of the trajectory | weight | reason |
|---|---|---|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
## Other decisions of 2026-10-06
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
- **Memory test:** waits until the data is near the size of the first SFT run.
## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
"""
open(os.path.join(ROOT, "docs", "stage2-build.md"), "w").write(doc)
print("written", len(doc))

320
train/build_stage2.py Normal file
View File

@@ -0,0 +1,320 @@
"""Stage 2 training set builder (Opus item C, 2026-10-06).
train/.venv/bin/python train/build_stage2.py [--out runs/stage2_data] [--max-tokens 48000] [--clas-cap 0.35]
[--min-per-kind 25] [--valid-frac 0.10] [--hook module:function]
Input: runs/traj (DeepSeek trajectories; the local Qwen series is never read). Steps: acceptance filter (score, end reason, no
harness text, loop repeats trimmed) -> eval overlap check -> Qwen 3.8 chat template (thinking off) with the tokenizer round trip
check -> samples over --max-tokens are dropped (never cut) -> CLAS capped at --clas-cap of the stage 2 samples (the rest goes to
`stage2_reserve.jsonl`) -> shortage report for kinds under --min-per-kind -> validation split by task family (K variants and all
trajectories of a task on the same side) -> optional hook (own-test score, weights) -> files + data card.
Loss mask: `assistant_spans` are character spans of the assistant turns (tool calls and the final report, with <|im_end|>);
nothing of the system turn (tool schemas), the user turn or the tool results is in a span. `assistant_tokens` counts the loss tokens.
"""
import argparse
import importlib
import re
import json
import os
import random
import statistics
import sys
import time
from collections import Counter, defaultdict
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
import to_qwen as tq # noqa: E402
from harness import mix, overlap # noqa: E402
from transformers import AutoTokenizer # noqa: E402
POOL = os.path.join(ROOT, "tasks_gen", "train")
START_PREFIX = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
LIST_TOOLS = ("sap_inactive_objects", "sap_search_object", "sap_usage_references")
READ_TOOLS = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
def foreign_name(name, own):
n = (name or "").upper()
m = START_PREFIX.match(n) or MID_PREFIX.match(n)
return bool(m) and m.group(1).upper() != own.upper()
def scrub_foreign(msgs, own_prefix):
"""Remove other runs' leftover objects from list results (the proxy hides them since 2026-10-06; the older records still have them).
Returns (messages, removed list entries, number of reads of a foreign object)."""
tool_of = {}
for m in msgs:
if m["role"] == "assistant":
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
tool_of[c.get("id")] = (c["function"]["name"], a)
removed = reads = 0
out = []
for m in msgs:
if m["role"] == "tool":
name, args = tool_of.get(m.get("tool_call_id"), ("", {}))
if name in READ_TOOLS and foreign_name(args.get("objectName"), own_prefix):
reads += 1
if name in LIST_TOOLS:
body = m["content"]
prefix = "ERROR: " if body.startswith("ERROR: ") else ""
try:
data = json.loads(body[len(prefix):])
except ValueError:
data = None
if isinstance(data, list):
keep = [d for d in data if not (isinstance(d, dict) and foreign_name(d.get("name"), own_prefix))]
if len(keep) != len(data):
removed += len(data) - len(keep)
m = dict(m, content=prefix + json.dumps(keep))
out.append(m)
return out, removed, reads
def pct(v, p):
v = sorted(v)
return v[min(len(v) - 1, int(p * len(v)))] if v else None
def family_of(task_id):
try:
t = json.load(open(os.path.join(POOL, task_id, "task.json")))
except OSError:
return task_id
return t.get("base_task") or task_id
def loss_tokens(tok, text, spans):
enc = tok(text, add_special_tokens=False, return_offsets_mapping=True)
n = 0
for (a, b) in enc["offset_mapping"]:
if any(a >= s and b <= e for s, e in spans):
n += 1
return len(enc["input_ids"]), n
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--traj", default=os.path.join(ROOT, "runs", "traj"))
ap.add_argument("--out", default=os.path.join(ROOT, "runs", "stage2_data"))
ap.add_argument("--max-tokens", type=int, default=48000)
ap.add_argument("--clas-cap", type=float, default=0.35)
ap.add_argument("--min-per-kind", type=int, default=25)
ap.add_argument("--min-score", type=float, default=80)
ap.add_argument("--valid-frac", type=float, default=0.10)
ap.add_argument("--seed", type=int, default=20261006)
ap.add_argument("--no-scrub", action="store_true", help="keep other runs' leftover objects in list results (default: removed)")
ap.add_argument("--keep-foreign-reads", action="store_true", help="keep trajectories in which the model read another run's object (default: dropped)")
ap.add_argument("--hook", help="module:function; function(row) -> None to drop, or a float weight (repeat factor), or a dict "
"{'keep': bool, 'weight': float, 'extra': {...}} (for example the own-test mutation score)")
a = ap.parse_args()
os.makedirs(a.out, exist_ok=True)
hook = None
if a.hook:
mod, fn = a.hook.split(":")
sys.path.insert(0, os.path.join(ROOT, "train"))
hook = getattr(importlib.import_module(mod), fn)
tok = AutoTokenizer.from_pretrained(os.path.expanduser("~/models/Qwen3.8-27B-4bit"))
evals = overlap.load_pool("eval")
report = {"built": time.strftime("%F %T"), "settings": vars(a), "dropped": defaultdict(Counter), "steps": {}}
# 1 accepted trajectories
rows = [json.loads(l) for l in open(os.path.join(a.traj, "summary.jsonl"))]
cand = []
for r in rows:
p = os.path.join(a.traj, r.get("run_dir") or "-", "record.json")
if not os.path.exists(p):
continue
rec = json.load(open(p))
ok, why = acc.judge(rec, r, a.min_score)
if not ok:
continue
msgs, removed = acc.trim_loops(rec["messages"])
scrubbed = freads = 0
if not a.no_scrub:
msgs, scrubbed, freads = scrub_foreign(msgs, rec["prefix"])
if freads and not a.keep_foreign_reads:
kind0 = mix.kind_of_task_dir(r["task"])
report["dropped"][kind0]["read_of_another_runs_object"] += 1
continue
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs), "scrubbed": scrubbed, "freads": freads})
report["steps"]["accepted_trajectories"] = len(cand)
# 2 eval overlap at build time (the task against every eval task)
kept = []
for c in cand:
tid = c["row"]["task"]
kind = mix.kind_of_task_dir(tid)
try:
hits = overlap.check(overlap.load_task(os.path.join(POOL, tid)), evals)
except (OSError, ValueError):
hits = []
if hits:
report["dropped"][kind]["eval_overlap"] += 1
continue
c["kind"] = kind
kept.append(c)
report["steps"]["after_eval_overlap"] = len(kept)
# 3 Qwen template + tokens + loss tokens; drop over the limit (never cut)
samples = []
for c in kept:
msgs = tq.to_template_messages(c["msgs"])
tools = c["rec"]["tools"]
text = tok.apply_chat_template(msgs, tools=tools, tokenize=False, enable_thinking=False)
prefix = tok.apply_chat_template(msgs[:2], tools=tools, tokenize=False, add_generation_prompt=True, enable_thinking=False)
if not text.startswith(prefix) or not tq.check_calls(msgs, text):
report["dropped"][c["kind"]]["template_roundtrip"] += 1
continue
spans = tq.spans(text)
n, nl = loss_tokens(tok, text, spans)
if n > a.max_tokens:
report["dropped"][c["kind"]]["over_%dk" % (a.max_tokens // 1000)] += 1
continue
t = c["rec"]["task"]
samples.append({"id": f"{c['row']['task']}_r{c['row']['run']}", "task": c["row"]["task"], "family": family_of(c["row"]["task"]),
"kind": c["kind"], "category": t.get("category"), "object_type": t.get("object_type"),
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"], "foreign_entries_removed": c["scrubbed"],
"score": c["rec"]["metadata"].get("score"), "teacher": c["rec"]["teacher"], "weight": 1.0,
"n_tokens": n, "assistant_tokens": nl, "text": text, "assistant_spans": spans})
report["steps"]["after_token_limit"] = len(samples)
# 4 CLAS cap: keep the CLAS samples with a repair and the highest score first; the rest is the reserve
by_kind = defaultdict(list)
for s in samples:
by_kind[s["kind"]].append(s)
non_clas = sum(len(v) for k, v in by_kind.items() if k != "CLAS")
clas = sorted(by_kind.get("CLAS", []), key=lambda s: (not s["repair"], -(s["score"] or 0), s["id"]))
limit = int(a.clas_cap / (1 - a.clas_cap) * non_clas) if non_clas else len(clas)
reserve = clas[limit:]
by_kind["CLAS"] = clas[:limit]
report["steps"]["clas_cap"] = {"clas_before": len(clas), "clas_kept": len(by_kind["CLAS"]), "clas_reserve": len(reserve), "limit": limit}
pool = [s for v in by_kind.values() for s in v]
# 5 shortage report (never filled by copies)
report["kinds"] = {}
for k in mix.TYPE_SHARE:
n = len(by_kind.get(k, []))
report["kinds"][k] = {"samples": n, "min_required": a.min_per_kind, "short": max(0, a.min_per_kind - n),
"families": len({s["family"] for s in by_kind.get(k, [])})}
# 6 hook (own-test score, weights)
if hook:
out = []
for s in pool:
r = hook(s)
if r is None or r is False:
report["dropped"][s["kind"]]["hook"] += 1
continue
if isinstance(r, dict):
if not r.get("keep", True):
report["dropped"][s["kind"]]["hook"] += 1
continue
s["weight"] = float(r.get("weight", 1.0))
s.update(r.get("extra") or {})
elif isinstance(r, (int, float)) and not isinstance(r, bool):
s["weight"] = float(r)
out.append(s)
pool = out
# 7 validation split by family, stratified by kind
rnd = random.Random(a.seed)
fam_kind = {}
for s in pool:
fam_kind.setdefault(s["family"], s["kind"])
valid_fam = set()
for k in mix.TYPE_SHARE:
fams = sorted(f for f, kk in fam_kind.items() if kk == k)
rnd.shuffle(fams)
nv = round(len(fams) * a.valid_frac)
if len(fams) >= 4:
nv = max(nv, 1)
valid_fam |= set(fams[:nv])
train = [s for s in pool if s["family"] not in valid_fam]
valid = [s for s in pool if s["family"] in valid_fam]
for name, data in (("stage2_train", train), ("stage2_valid", valid), ("stage2_reserve", reserve)):
with open(os.path.join(a.out, name + ".jsonl"), "w") as f:
for s in data:
f.write(json.dumps(s) + "\n")
# 8 report and data card
def summary(data):
toks = [s["n_tokens"] for s in data]
return {"samples": len(data), "tokens": sum(toks), "loss_tokens": sum(s["assistant_tokens"] for s in data),
"p50": pct(toks, .5), "p90": pct(toks, .9), "p95": pct(toks, .95), "max": max(toks or [0]),
"by_kind": dict(Counter(s["kind"] for s in data)), "by_category": dict(Counter(s["category"] for s in data)),
"repair_share": round(sum(s["repair"] for s in data) / max(len(data), 1), 2)}
report["train"], report["valid"], report["reserve"] = summary(train), summary(valid), summary(reserve)
report["dropped"] = {k: dict(v) for k, v in report["dropped"].items()}
s1 = json.load(open(os.path.join(ROOT, "train", "data", "stats.json"))) if os.path.exists(os.path.join(ROOT, "train", "data", "stats.json")) else {}
report["stage1_stats"] = s1
json.dump(report, open(os.path.join(a.out, "build_report.json"), "w"), indent=1, default=str)
open(os.path.join(a.out, "README.md"), "w").write(data_card(report))
print(json.dumps({k: report[k] for k in ("steps", "train", "valid", "reserve", "dropped")}, indent=1, default=str)[:3500])
print("short kinds:", {k: v["short"] for k, v in report["kinds"].items() if v["short"]})
def data_card(r):
t, v = r["train"], r["valid"]
kinds = "\n".join(f"| {k} | {x['samples']} | {x['families']} | {x['short'] or ''} |" for k, x in r["kinds"].items())
drop = "\n".join(f"- {k}: {d}" for k, d in r["dropped"].items()) or "- nothing dropped"
return f"""---
license: mit
task_categories: [text-generation]
tags: [abap, sap, agentic, tool-use, sft]
private: true
---
# ABAP stage 2 agent trajectories (built {r['built']})
Tool-using ABAP development trajectories for supervised fine-tuning of Qwen 3.8 27B. A teacher model solved generated ABAP tasks on a real
SAP ABAP Platform system (A4H, SAP_BASIS 816) through ADT tools; only runs that passed the harness gates and scored at least {r['settings']['min_score']:.0f} of 100 are kept.
## Sources and licenses
- Teacher: DeepSeek V4.1 Flash (MIT license), via Ollama cloud. No output of Claude or other restricted models.
- Tasks: generated by the same teacher (spec, seed objects, hidden ABAP Unit tests, reference), validated on A4H (oracle 100, null 0, mutation check). Eval tasks are not in this data (overlap check at build time).
- The local Qwen runs (series A) are not in this data.
## Format
One JSON per line: `text` (the whole conversation in the Qwen 3.8 chat template, thinking off, Qwen XML tool call format, the 20 tool schemas in the system turn),
`assistant_spans` (character spans that carry the loss: assistant turns with tool calls and the final report; none of the system turn, user turn or tool results),
`n_tokens`, `assistant_tokens`, `kind`, `category`, `object_type`, `family` (task family: a K variant has the family of its base task), `repair` (an error followed by a fix), `score`, `weight`.
## Size
| split | samples | tokens | loss tokens | p50 | p90 | p95 | max | repair share |
|---|---|---|---|---|---|---|---|---|
| train | {t['samples']} | {t['tokens']} | {t['loss_tokens']} | {t['p50']} | {t['p90']} | {t['p95']} | {t['max']} | {t['repair_share']} |
| valid | {v['samples']} | {v['tokens']} | {v['loss_tokens']} | {v['p50']} | {v['p90']} | {v['p95']} | {v['max']} | {v['repair_share']} |
Samples over {r['settings']['max_tokens']} tokens are dropped, never cut. CLAS is capped at {int(100 * r['settings']['clas_cap'])} % of the samples (the rest is in `stage2_reserve.jsonl`).
## Kinds (target share: CLAS 28, INTF 7, CDS 25, FUNC 15, PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5 percent)
| kind | samples | families | short of the minimum {r['settings']['min_per_kind']} |
|---|---|---|---|
{kinds}
## Dropped
{drop}
## Known limits
- Small data (see the table); kinds marked short have fewer than the minimum samples.
- Until 2026-10-06 the test system still held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes). Their names are removed from list results (`sap_inactive_objects`, searches) in these samples; trajectories in which the model read such an object are dropped.
- The proxy added local abaplint messages (`syntaxCheck`) to some failed writes; the EPOD server does not do this itself.
- The teacher sees only the tool results of its own runs; a run that repaired an error is kept with its error (60 to 65 % of the samples).
- Validation split is by task family; do not mix valid samples into the training set.
"""
if __name__ == "__main__":
main()

79
train/foreign_scan.py Normal file
View File

@@ -0,0 +1,79 @@
"""Read-only scan: which runs saw other runs' objects in tool results, or read them (teardown/proxy bug of 2026-10-06).
python3 train/foreign_scan.py -> prints a summary and writes runs/analysis/foreign_scan.json
"""
import glob
import json
import os
import re
import sys
from collections import Counter
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
from harness.task import prefix_for # noqa: E402
START = re.compile(r"Z\d[0-9A-Z]{6}_")
MID_NAME = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
GROUPS = [("official baseline Qwen (MacBook, 11 tasks)", "runs/stage1/baseline_qwen/*"),
("Devstral baseline (dropped)", "runs/archive/devstral/baseline_devstral/*"),
("older Qwen baselines (archive)", "runs/archive/qwen38/*/*"),
("smoke", "runs/stage1/smoke/*"),
("eval empirical filter (DeepSeek on eval candidates)", "runs/emp/*"),
("eval generation validation (oracle, null, mutants)", "runs/gen/*"),
("early pilot runs", "runs/2*_*"),
("training trajectories (DeepSeek)", "runs/traj/*_G*"),
("series A (local Qwen)", "runs/local_qwen/runs/*")]
def own_prefix(d):
m = re.match(r"^(\d+)_([GT]\d+)_", os.path.basename(d))
return prefix_for(int(m.group(1)), m.group(2)).upper() if m else None
def scan(d):
own = own_prefix(d)
p = os.path.join(d, "trajectory.jsonl")
if not own or not os.path.exists(p):
return None
seen, reads, tools = set(), 0, Counter()
for l in open(p):
try:
e = json.loads(l)
except ValueError:
continue
if "tool" not in e:
continue
text = (e.get("result") or "").upper()
found = {x for x in START.findall(text) if x != own and x[1].isdigit()}
found |= {m.group(1).upper() for m in re.finditer(r"\b[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", text) if m.group(1).upper() != own}
if found:
seen |= found
tools[e["tool"]] += 1
name = str((e.get("args") or {}).get("objectName", "")).upper()
m = START.match(name) or MID_NAME.match(name)
pre = (m.group(1) if m and m.re is MID_NAME else (m.group(0) if m else "")).upper()
if e["tool"] in READ and pre and pre != own:
reads += 1
return {"run": os.path.basename(d), "foreign_prefixes_seen": len(seen), "foreign_reads": reads, "tools": dict(tools)}
def main():
out = {}
for label, pat in GROUPS:
dirs = [d for d in sorted(glob.glob(os.path.join(ROOT, pat))) if os.path.isdir(d)]
rows = [r for r in (scan(d) for d in dirs) if r]
seen = [r for r in rows if r["foreign_prefixes_seen"]]
reads = [r for r in rows if r["foreign_reads"]]
out[label] = {"runs": len(rows), "runs_with_foreign_names": len(seen), "runs_with_foreign_reads": len(reads),
"tools": dict(sum((Counter(r["tools"]) for r in seen), Counter())),
"affected": [r["run"] for r in seen][:40], "read_runs": [r["run"] for r in reads]}
print(f"{label}: {len(rows)} runs | foreign names in tool results: {len(seen)} | foreign reads: {len(reads)} | tools {out[label]['tools']}")
os.makedirs(os.path.join(ROOT, "runs", "analysis"), exist_ok=True)
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "foreign_scan.json"), "w"), indent=1)
if __name__ == "__main__":
main()

182
train/hf_train_bf16.py Normal file
View File

@@ -0,0 +1,182 @@
# /// script
# requires-python = ">=3.10"
# dependencies = ["unsloth", "datasets", "transformers", "huggingface_hub"]
# ///
"""Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go.
Memory test first (a few dollars, finds the largest sequence length that fits, no data needed):
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000,64000
Real run (after the memory test and Kral's decision on the ratio):
hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s2-epochs 3 --s2-loss-share 0.6 --out erhankeseli/abap-mixed-adapter
Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns;
none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut.
Base model in bf16 (no nf4): the adapter then fits the MLX 4-bit base better (stage 1 finding, train/STATE.md).
"""
import argparse
import json
import os
import random
import time
def tokenize_masked(tok, text, spans, max_len):
"""input_ids and labels; labels are -100 outside the spans (spans=None: loss on every token). None when longer than max_len."""
enc = tok(text, add_special_tokens=False, return_offsets_mapping=True)
ids = enc["input_ids"]
if len(ids) > max_len:
return None
if spans is None:
return ids, list(ids)
labels = []
for tid, (a, b) in zip(ids, enc["offset_mapping"]):
labels.append(tid if any(a >= s and b <= e for s, e in spans) else -100)
return ids, labels
def s1_epochs_for(share, s2_loss_tokens, s2_epochs, s1_tokens):
"""Epochs of stage 1 so that stage 2 carries `share` of all loss tokens (Kral + Opus 2026-10-06: share = 0.6)."""
s2_total = s2_loss_tokens * s2_epochs
return (1 - share) / share * s2_total / max(s1_tokens, 1)
def copies(weight, rnd):
"""Weight 1 = one copy per epoch, 2 = two, 0.5 = a copy in half of the epochs (a down-weight, not a drop)."""
w = max(float(weight), 0.0)
return int(w) + (1 if rnd.random() < w - int(w) else 0)
def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
rnd = random.Random(seed)
rows = []
for ep in range(int(s1_epochs)):
rows += [("s1", r["text"], None, 1.0) for r in stage1]
frac = s1_epochs - int(s1_epochs)
if frac:
rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))]
for ep in range(int(s2_epochs)):
for r in stage2:
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * copies(r.get("weight", 1.0), rnd)
rnd.shuffle(rows)
out, skipped = [], {"s1": 0, "s2": 0}
for src, text, spans, _ in rows:
t = tokenize_masked(tok, text, spans, max_len)
if t is None:
skipped[src] += 1
continue
out.append({"input_ids": t[0], "labels": t[1], "src": src})
return out, skipped
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--model", default="Qwen/Qwen3.8-27B")
ap.add_argument("--stage1", default="erhankeseli/abap-stage1-data")
ap.add_argument("--stage2", default="erhankeseli/abap-stage2-data")
ap.add_argument("--s1-file", default="train.jsonl")
ap.add_argument("--s2-file", default="stage2_train.jsonl")
ap.add_argument("--s2-valid", default="stage2_valid.jsonl")
ap.add_argument("--s1-epochs", type=float, default=None, help="epochs of stage 1; default: computed from --s2-loss-share")
ap.add_argument("--s2-epochs", type=float, default=3.0)
ap.add_argument("--s2-loss-share", type=float, default=0.6, help="share of the loss tokens that stage 2 carries (decision 2026-10-06: 0.6)")
ap.add_argument("--max-seq", type=int, default=48000)
ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--alpha", type=int, default=32)
ap.add_argument("--lr", type=float, default=5e-5)
ap.add_argument("--grad-accum", type=int, default=4)
ap.add_argument("--warmup-steps", type=int, default=20)
ap.add_argument("--no-offload", action="store_true", help="gradient checkpointing on the GPU instead of the Unsloth offload")
ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter")
ap.add_argument("--save-every", type=int, default=0)
ap.add_argument("--memory-test", action="store_true")
ap.add_argument("--sweep", default="16000,32000,48000,64000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length")
a = ap.parse_args()
import torch
from unsloth import FastLanguageModel
from huggingface_hub import HfApi, hf_hub_download
token = os.environ["HF_TOKEN"]
seq_for_model = max([a.max_seq] + ([int(x) for x in a.sweep.split(",")] if a.memory_test else []))
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=seq_for_model, load_in_4bit=False, dtype=torch.bfloat16, token=token)
model = FastLanguageModel.get_peft_model(
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"],
use_gradient_checkpointing=True if a.no_offload else "unsloth", random_state=20261006)
print("GPU", torch.cuda.get_device_name(0), "x", torch.cuda.device_count(), "| weights allocated GB", round(torch.cuda.memory_allocated() / 2**30, 1), flush=True)
if a.memory_test:
text = "CLASS zcl_demo DEFINITION PUBLIC. METHODS run IMPORTING iv TYPE i. ENDCLASS.\n" * 4000
ids_all = tok(text, add_special_tokens=False)["input_ids"]
opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=1e-5)
model.train()
result = {"gpu": torch.cuda.get_device_name(0), "gpus": torch.cuda.device_count(), "offload": not a.no_offload, "sweep": []}
for n in [int(x) for x in a.sweep.split(",")]:
ids = torch.tensor([(ids_all * (n // len(ids_all) + 1))[:n]], device="cuda")
labels = ids.clone()
labels[:, : int(n * 0.68)] = -100 # about 32 % loss tokens, as in the stage 2 data
torch.cuda.reset_peak_memory_stats()
try:
times = []
for _ in range(a.steps):
t0 = time.time()
loss = model(input_ids=ids, labels=labels).loss
loss.backward()
opt.step()
opt.zero_grad(set_to_none=True)
torch.cuda.synchronize()
times.append(round(time.time() - t0, 1))
row = {"seq_len": n, "ok": True, "peak_gb": round(torch.cuda.max_memory_allocated() / 2**30, 1),
"reserved_gb": round(torch.cuda.max_memory_reserved() / 2**30, 1), "step_seconds": times, "loss": float(loss)}
except torch.cuda.OutOfMemoryError as e:
row = {"seq_len": n, "ok": False, "error": "OOM", "peak_gb": round(torch.cuda.max_memory_allocated() / 2**30, 1)}
result["sweep"].append(row)
print("MEMTEST", json.dumps(row), flush=True)
break
result["sweep"].append(row)
print("MEMTEST", json.dumps(row), flush=True)
torch.cuda.empty_cache()
print("RESULT", json.dumps(result), flush=True)
HfApi(token=token).upload_file(path_or_fileobj=json.dumps(result, indent=1).encode(), path_in_repo="memtest_%s_%dx.json" % (
result["gpu"].replace(" ", "_"), result["gpus"]), repo_id=a.out, repo_type="model", create_pr=False) if a.out else None
return
def rows(repo, fname):
p = hf_hub_download(repo, fname, repo_type="dataset", token=token)
return [json.loads(l) for l in open(p)]
s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file)
if a.s1_epochs is None:
s1_tok = sum(len(tok(r["text"], add_special_tokens=False)["input_ids"]) for r in s1)
s2_loss = sum(r["assistant_tokens"] * float(r.get("weight", 1.0)) for r in s2)
a.s1_epochs = s1_epochs_for(a.s2_loss_share, s2_loss, a.s2_epochs, s1_tok)
print("STAGE1 EPOCHS", round(a.s1_epochs, 3), "(stage 1 tokens", s1_tok, ", stage 2 loss tokens per epoch", round(s2_loss), ", share", a.s2_loss_share, ")", flush=True)
ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006)
loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex)
print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens",
round(sum(sum(1 for x in e["labels"] if x != -100) for e in ex if e["src"] == "s2") / max(loss_tok, 1), 2), flush=True)
from datasets import Dataset
from transformers import Trainer, TrainingArguments
class Collate:
def __call__(self, batch): # batch size 1; no padding needed
b = batch[0]
return {"input_ids": torch.tensor([b["input_ids"]]), "labels": torch.tensor([b["labels"]]),
"attention_mask": torch.ones(1, len(b["input_ids"]), dtype=torch.long)}
args = TrainingArguments(output_dir="out", per_device_train_batch_size=1, gradient_accumulation_steps=a.grad_accum, num_train_epochs=1,
learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=a.warmup_steps, optim="adamw_8bit", bf16=True,
logging_steps=1, save_strategy="no", report_to="none", seed=20261006, remove_unused_columns=False)
tr = Trainer(model=model, args=args, train_dataset=Dataset.from_list(ex), data_collator=Collate())
t0 = time.time()
tr.train()
info = {"examples": len(ex), "skipped": skipped, "train_seconds": round(time.time() - t0, 1), "rank": a.rank, "alpha": a.alpha, "lr": a.lr,
"peak_gpu_gb": round(torch.cuda.max_memory_allocated() / 2**30, 1), "s1_epochs": a.s1_epochs, "s2_epochs": a.s2_epochs}
print("RESULT", json.dumps(info), flush=True)
model.save_pretrained("adapter")
json.dump(info, open("adapter/job_result.json", "w"))
HfApi(token=token).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model")
print("PUSHED", a.out)
if __name__ == "__main__":
main()

58
train/hooks_example.py Normal file
View File

@@ -0,0 +1,58 @@
"""Hooks for train/build_stage2.py (--hook hooks_example:own_test_weight). A hook gets one sample (a dict) and returns
None / False (drop), a number (repeat weight), or {"keep": bool, "weight": float, "extra": {...}}."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
def identity(row):
return 1.0
# Weights from the own-test mutation score (Kral + Opus 2026-10-06: down-weight, do not drop). Weight w means: w copies per epoch, a fraction is a
# copy in that share of the epochs (0.5 = every second epoch on average). The acceptance filter itself is unchanged.
OWN_TEST_WEIGHTS = {
"reliable, score >= 0.75": 1.0, # the own tests pass on the correct reference and kill at least 3 of 4 mutants
"reliable, score 0.5 to 0.75": 0.75,
"reliable, score < 0.5": 0.5,
"unreliable (own tests fail on the correct reference)": 0.5, # they may encode model specific behavior
"no own tests": 0.5, # accepted at 85 points at most; the habit of writing tests is part of the behavior we want
"no signal (PROG, no mutants, not scored)": 1.0,
}
def own_test_class(m):
if m is None:
return "no signal (PROG, no mutants, not scored)"
st = m.get("status")
if st == "no_own_tests":
return "no own tests"
if st != "scored":
return "no signal (PROG, no mutants, not scored)"
if not m.get("tests_pass_on_reference"):
return "unreliable (own tests fail on the correct reference)"
s = m.get("score")
if s is None:
return "no signal (PROG, no mutants, not scored)"
return "reliable, score >= 0.75" if s >= 0.75 else "reliable, score 0.5 to 0.75" if s >= 0.5 else "reliable, score < 0.5"
def own_test_weight(row):
"""Item D: reads runs/traj/<run>/own_test_mutation.json, stores the score as extra data and sets the weight by OWN_TEST_WEIGHTS."""
run = row["id"].split("_r")[-1]
m = None
for d in os.listdir(os.path.join(ROOT, "runs", "traj")):
if d.startswith(run + "_"):
p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json")
if os.path.exists(p):
m = json.load(open(p))
break
cls = own_test_class(m)
return {"keep": True, "weight": OWN_TEST_WEIGHTS[cls],
"extra": {"own_test_class": cls, "own_test_mutation": (m or {}).get("score"), "own_test_reliable": bool((m or {}).get("tests_pass_on_reference"))}}
def repair_up(row):
"""Example of a weight: a trajectory with a repair counts twice."""
return 2.0 if row.get("repair") else 1.0

71
train/mem_table.py Normal file
View File

@@ -0,0 +1,71 @@
"""Writes docs/bf16-memory.md (estimates for the bf16 LoRA run of Qwen 3.8 27B by GPU and sequence length)."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
H, L, V, I = 5120, 64, 248320, 17408
r = 16
lora = 16 * r * ((H + 6144) + (H + 1024) + (H + 1024) + (6144 + H)) + 48 * r * ((H + 10240) + (H + 6144) + (6144 + H)) + L * r * ((H + I) * 2 + (I + H))
W = 27.8e9 * 2 / 2**30
lora_gb = lora * 10 / 2**30
def est(seq, offload, chunked):
ckpt = 0.0 if offload else seq * H * 2 * L / 2**30
layer = seq * I * 2 * 4 / 2**30 + 3
logits = (seq * V * 2 * 3 / 2**30) if not chunked else 3.0
return W + lora_gb + ckpt + layer + logits, dict(weights=W, lora=lora_gb, ckpt=ckpt, layer=layer, logits=logits)
gpus = [("a100-large", "1x A100 80 GB", 80), ("rtx-pro-6000", "1x RTX PRO 6000 96 GB", 96), ("h200", "1x H200 141 GB", 141), ("h200x2", "2x H200 282 GB (model parallel)", 282)]
tab = "| seq | checkpoints | loss | estimated peak GB | " + " | ".join(g[1] for g in gpus) + " |\n|---|---|---|---|" + "---|" * len(gpus) + "\n"
for seq in (16000, 32000, 48000, 64000):
for off in (True, False):
for ch in (True, False):
t, _ = est(seq, off, ch)
cells = ["fits" if t <= g[2] * 0.92 else ("tight" if t <= g[2] else "no") for g in gpus]
tab += f"| {seq // 1000}k | {'CPU offload (Unsloth)' if off else 'on the GPU'} | {'chunked' if ch else 'full logits'} | {t:.0f} | " + " | ".join(cells) + " |\n"
_, b48 = est(48000, True, True)
_, b64 = est(64000, True, True)
doc = f"""# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
## Model (from config.json of Qwen3.8-27B)
64 layers (48 linear attention, 16 full attention), hidden {H}, MLP {I}, vocab {V}, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **{lora / 1e6:.0f} M trainable parameters**.
## Memory by component (GB; estimate, not a measurement)
| component | 48k | 64k | how |
|---|---|---|---|
| weights bf16 | {b48['weights']:.1f} | {b64['weights']:.1f} | 27.8 B x 2 bytes |
| LoRA weights + grads + 8-bit Adam | {b48['lora']:.1f} | {b64['lora']:.1f} | {lora / 1e6:.0f} M x 10 bytes |
| layer inputs for the backward pass | {48000 * H * 2 * L / 2**30:.1f} on the GPU, about 0 with the Unsloth CPU offload (then {48000 * H * 2 * L / 2**30:.0f} GB host RAM) | {64000 * H * 2 * L / 2**30:.1f} / about 0 (host RAM {64000 * H * 2 * L / 2**30:.0f} GB) | seq x 5120 x 2 bytes x 64 layers |
| recompute peak of one layer | {b48['layer']:.1f} | {b64['layer']:.1f} | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
| logits and loss | {b48['logits']:.0f} chunked, {48000 * V * 2 * 3 / 2**30:.0f} full logits | {b64['logits']:.0f} chunked, {64000 * V * 2 * 3 / 2**30:.0f} full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
## Fit by GPU (92 % of the card counted as usable)
{tab}
Reading:
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about {b48['weights'] + b48['lora'] + b48['layer'] + b48['logits']:.0f} GB at 48k, {b64['weights'] + b64['lora'] + b64['layer'] + b64['logits']:.0f} GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
## Order when the test is run
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
## Time and cost (estimate from the nf4 run: {6750 / 30.5:.0f} tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
A100 about {3.3e6 / 300 / 3600:.1f} h (about {3.3e6 / 300 / 3600 * 2.5:.0f} USD), H200 about {3.3e6 / 650 / 3600:.1f} h (about {3.3e6 / 650 / 3600 * 5:.0f} USD).
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
H200 about {10e6 / 650 / 3600:.1f} h (about {10e6 / 650 / 3600 * 5:.0f} USD), A100 about {10e6 / 300 / 3600:.0f} h (about {10e6 / 300 / 3600 * 2.5:.0f} USD). The real number comes from the memory test (`step_seconds`).
"""
open(os.path.join(ROOT, "docs", "bf16-memory.md"), "w").write(doc)
print("written")

42
train/own_test_report.py Normal file
View File

@@ -0,0 +1,42 @@
"""Summary of the own-test mutation scores (item D): docs/own-test-mutation.md. Run after `python3 -m harness.owntests`."""
import glob
import json
import os
import statistics
from collections import Counter, defaultdict
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
res = [json.load(open(p)) for p in glob.glob(os.path.join(ROOT, "runs", "traj", "*", "own_test_mutation.json"))]
by = Counter(r.get("status") for r in res)
scored = [r for r in res if r.get("status") == "scored"]
ref_ok = [r for r in scored if r.get("tests_pass_on_reference")]
per_kind = defaultdict(list)
for r in ref_ok:
if r.get("score") is not None:
per_kind[r["kind"]].append(r["score"])
sc = [r["score"] for r in ref_ok if r.get("score") is not None]
bins = Counter("1.0" if s == 1 else ">=0.75" if s >= 0.75 else ">=0.5" if s >= 0.5 else "<0.5" for s in sc)
low = sorted([r for r in ref_ok if r.get("score") is not None and r["score"] < 0.75], key=lambda r: r["score"])
doc = f"""# Own-test mutation score of the accepted trajectories (item D, metadata only)
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
## Result
{len(res)} accepted trajectories looked at: {dict(by)}.
- Scored: {len(scored)}; the model's tests also passed on the correct reference: {len(ref_ok)} (the others are not reliable: a test that fails on the correct solution kills every mutant).
- Mean score over the reliable ones: {statistics.mean(sc):.2f}; median {statistics.median(sc):.2f}; distribution {dict(bins)} (n = {len(sc)}).
- By kind (mean, n): {{{", ".join(f"{k}: {statistics.mean(v):.2f} ({len(v)})" for k, v in sorted(per_kind.items()))}}}
- `no_own_tests`: {by.get('no_own_tests', 0)} trajectories were accepted without any own test (at most 85 points); `no_mutants`: {by.get('no_mutants', 0)}; `not_supported` (PROG: tests are inside the program): {by.get('not_supported', 0)}.
## Weakest (score under 0.75, tests pass on the reference)
""" + "\n".join(f"- {r['task']} run {r['run']} {r['kind']}: score {r['score']} ({r['killed']} of {r['valid']} valid mutants killed)" for r in low[:15]) + """
## Use
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).
"""
open(os.path.join(ROOT, "docs", "own-test-mutation.md"), "w").write(doc)
print(doc[:1500])

75
train/requeue_foreign.py Normal file
View File

@@ -0,0 +1,75 @@
"""Trajectories in which the model read another run's leftover object go back to the pending pool (Kral + Opus 2026-10-06).
python3 train/requeue_foreign.py dry run
python3 train/requeue_foreign.py --apply moves their rows from runs/traj/summary.jsonl to runs/traj/summary_excluded.jsonl
(backup: summary.jsonl.bak-<time>); the task then counts as not run and is run again
after the reset in the normal order (kind deficit). The run folders stay.
"""
import json
import os
import shutil
import sys
import time
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
from harness import mix # noqa: E402
START = __import__("re").compile(r"^(Z\d[0-9A-Z]{6}_)", __import__("re").I)
MID = __import__("re").compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", __import__("re").I)
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
def foreign_reads(rec):
own = rec["prefix"].upper()
n = 0
for m in rec["messages"]:
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
if c["function"]["name"] not in READ:
continue
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
name = str(a.get("objectName", "")).upper()
mm = START.match(name) or MID.match(name)
if mm and mm.group(1).upper() != own:
n += 1
return n
def main():
apply = "--apply" in sys.argv
path = os.path.join(ROOT, "runs", "traj", "summary.jsonl")
rows = [json.loads(l) for l in open(path)]
keep, out = [], []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
if os.path.exists(p):
rec = json.load(open(p))
if acc.judge(rec, r, 80)[0]:
n = foreign_reads(rec)
if n:
out.append(dict(r, excluded_reason="read of another run's object", foreign_reads=n, kind=mix.kind_of_task_dir(r["task"]),
excluded_at=time.strftime("%F %T")))
continue
keep.append(r)
print(len(out), "accepted trajectories with a foreign read:", [(o["task"], o["attempt"], o["kind"], o["foreign_reads"]) for o in out])
if not apply:
return
shutil.copy(path, path + ".bak-" + time.strftime("%Y%m%d-%H%M%S"))
with open(os.path.join(ROOT, "runs", "traj", "summary_excluded.jsonl"), "a") as f:
for o in out:
f.write(json.dumps(o) + "\n")
open(path, "w").write("".join(json.dumps(r) + "\n" for r in keep))
print("moved; summary.jsonl now", len(keep), "rows")
if __name__ == "__main__":
main()