Compare commits

..

7 Commits

28 changed files with 1958 additions and 23 deletions

48
docs/bf16-memory.md Normal file
View File

@@ -0,0 +1,48 @@
# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
**No GPU job is started without Kral's go.** Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000` is the memory test, no data needed, 3 optimizer steps per length, stops at the first OOM, uploads the result as `memtest_*.json` to the output repo).
## Model (from config.json of Qwen3.8-27B)
64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**.
## Memory by component at 48k tokens (GB; the estimate, not a measurement)
| component | GB | how |
|---|---|---|
| weights bf16 | 51.8 | 27.8 B x 2 bytes |
| LoRA weights + grads + 8-bit Adam | 1.0 | 107 M x 10 bytes |
| layer inputs for the backward pass (checkpoints) | 29.3 on the GPU, about 0 with the Unsloth CPU offload | 48k x 5120 x 2 bytes x 64 layers; the offload needs 29 GB of host RAM (all flavors have 142 GB or more) |
| recompute peak of one layer | 9.2 | MLP tensors 48k x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
| logits and loss | 3 chunked, 67 with full logits | full logits: 48k x 248320 x (bf16 + fp32 upcast + grad); only a chunked or fused cross entropy is realistic |
## Fit by GPU (92 % of the card counted as usable)
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (needs model parallel) |
|---|---|---|---|---|---|---|---|
| 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits |
| 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits |
| 16k | on the GPU | chunked | 71 | fits | fits | fits | fits |
| 16k | on the GPU | full logits | 90 | no | tight | fits | fits |
| 32k | CPU offload (Unsloth) | chunked | 63 | fits | fits | fits | fits |
| 32k | CPU offload (Unsloth) | full logits | 104 | no | no | fits | fits |
| 32k | on the GPU | chunked | 82 | no | fits | fits | fits |
| 32k | on the GPU | full logits | 124 | no | no | fits | fits |
| 48k | CPU offload (Unsloth) | chunked | 65 | fits | fits | fits | fits |
| 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits |
| 48k | on the GPU | chunked | 94 | no | tight | fits | fits |
| 48k | on the GPU | full logits | 158 | no | no | no | fits |
Reading:
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, with the CPU offload) fits 32k and 48k; the A100 and the RTX PRO 6000 do not. The run needs a fused or chunked cross entropy
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is not known). The memory test shows it at once: if `seq_len 16000` already fails on an H200, the loss is the cause.
Fallback (not built yet, small): compute the hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
- **A100 80 GB (2.50 USD/hour) is only possible with the CPU offload and a chunked loss** (about 65 GB estimated at 48k: no margin). H200 141 GB (5 USD/hour) fits with margin;
RTX PRO 6000 96 GB (2.75 USD/hour) fits with the offload and the chunked loss (est. 65 GB).
- Samples over 32k are a minority (p90 39560 tokens in the current data, p50 24.8k): if 48k does not fit, 32k would drop 15 of 52 samples (29 %; 10 of them DDLS), so test 48k first.
- Multi-GPU (`a100x4`, `h200x2`): possible only with model parallelism (`device_map`), one card works at a time, Unsloth multi-GPU is limited; not recommended before the single card test.
## Proposed order
1. Memory test on **h200** (about 30 minutes, 2.5 USD): sweep 16000,32000,48000 with the offload. If 48k fits with margin: stay on H200 or try rtx-pro-6000 for the real run.
2. If the loss is the problem: build the chunked loss (about 2 hours of work), repeat the test.
3. Real run after the stage 1 : stage 2 ratio decision (below).
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 is faster, H200 about 2 to 2.5 times an A100)
Tokens of the proposed run (1 epoch stage 1 and 3 epochs of the 47 stage 2 samples as built today: 2.53 M + 3 x 1.28 M = 6.4 M tokens): A100 about 5.9 h (about 15 USD), H200 about 2.7 h (about 14 USD). The real number comes from the memory test (`step_seconds`).

View File

@@ -0,0 +1,55 @@
# Stage 2 data analysis (accepted DeepSeek trajectories), 2026-10-06
Script `train/analyze.py`, numbers in `runs/analysis/analysis.json`. Base: 90 accepted trajectories of 119 runs (acceptance 76 %).
## 1. Repair taxonomy: are the three step-3 errors covered?
133 failed writes occurred inside accepted trajectories; each of the three common errors appears and is repaired:
| error (roadmap step 3) | failed writes | trajectories | repaired in the same trajectory | example |
|---|---|---|---|---|
| `TYPE c LENGTH n` in a method signature | 10 | 3 | 3 | Unable to interpret "4". Possible causes of error include incorrect spellings or comma errors. |
| reserved word as a field or parameter name | 18 | 16 | 16 | Field "VALUE" is unknown. |
| name longer than 30 characters | 25 | 23 | 23 | The name "SORTS_BY_TOTAL_DESC_THEN_VARIETY" is longer than the allowed 30 characters. |
Reading: the long names (23 trajectories) and the reserved words (16, by a text match: the label is generous, it also counts "Field VALUE is unknown") are well covered.
**`TYPE c LENGTH` in a signature is thin: 3 trajectories** (10 failed writes, all repaired). The SAP message is only "save operation failed" or, with the proxy hint,
`Statement does not exist ... METHODS`. This is the error that the proxy syntaxCheck exists for; after the reset the error-targeted slots ("named-type") should be
run first for CLAS and FUNC to get more of these. Other frequent errors (top list in the json): the testclasses include is not reachable with `sap_push_element` (9 times),
"save operation failed" without detail (30 with the hint), "Field X is unknown", "Test method can be defined only in test classes".
## 2. Behaviors that Qwen lacks (series A, all 20 runs) versus the teacher (accepted trajectories)
| behavior | DeepSeek (accepted) | Qwen series A |
|---|---|---|
| runs with at least one failed write | 60 of 90 | 14 of 20 |
| **repair after the first error** (a different source that is written successfully) | **59 of 60 (98 %)** | **6 of 14 (43 %)** |
| median calls from the first error to the repair | 1 | 2.0 |
| searches/reads before the first write (median / p90) | 3.0 / 19 | 4.0 / 15 |
| longest series of `sap_search_object` calls | 7 | **61** |
| runs with the same source pushed again | 2 of 90 | **11 of 20** |
| runs without any write | 4 | 4 |
The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61.
These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories).
Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection.
For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (29 runs), not analysed here.
## 3. Near duplicates
Pairwise comparison of the code that each accepted trajectory wrote (5-word shingles, prefix removed) and of the final reports: **no pair above 0.8 (code) or 0.85 (report)**, also none
between the two trajectories of one task. The data is not repetitive at this level.
## 4. empty_response (13 of 119 runs, 11 %)
By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and
contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call;
in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862,
the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 10 longer than 20000, 22 longer than 16000.
Fix: `docs/empty-response.md`.
## 5. What it means for the data
- Keep the repair behavior and the "write soon" behavior: they are present. Weight the error-targeted slots (named-type) up after the reset.
- Do not rely on `TYPE c LENGTH` repairs: only 3 examples.
- The DDLS share of empty runs (6 of 13) explains much of the CDS rejection.

23
docs/empty-response.md Normal file
View File

@@ -0,0 +1,23 @@
# empty_response fix (2026-10-06)
**Problem.** 13 of 119 DeepSeek runs (11 %) ended with `empty_response`: one turn used the whole output limit (32000 tokens) for reasoning,
with no content and no tool call, and the retry (same prompt, same temperature, up to two) did it again in 7 of 13 runs. Each such run costs up to
3 x 32000 output tokens and is lost (DDLS 6, PROG 3, CLAS 2, TABL 2). Analysis: `docs/data-analysis-stage2.md`, section 4.
**What the data says.** The empty turn comes after a small read result (50 to 6000 characters), in contexts of 13k to 63k tokens, not after a large result. Turns
over all runs: p50 1.1k, p90 7.9k tokens; legitimate long turns (large test classes) reach 30k; only 3 of 1997 turns longer than 24k were legitimate.
**Implemented (harness side only, `harness/agents.py`, `harness/trajectories.py`)**
1. Output cap per turn **24000** (was 32000; saves a quarter of every runaway turn, costs 3 of 1997 legitimate turns).
2. **One retry at temperature 0.8** (was two retries at 0.2): another sample instead of the same runaway.
3. **Stream guard** (off until verified): the request is streamed; when a turn has produced only reasoning for `STREAM_GUARD` tokens (suggested 9000) and no content and
no tool call, the stream is cut and the turn counts as empty (retry as in 2). Saves most of the cost of a runaway turn. Tested with a fake streaming server (text, tool call,
runaway: cut after 1500 estimated tokens in 0.01 s). **Not tested on the cloud model**: it needs the Ollama cloud stream to carry the reasoning in `delta.reasoning`
(or `reasoning_content` / `thinking`) and tool calls as streamed deltas; if the stream cannot be parsed the agent falls back to the normal request. Usage of a cut turn is estimated
(characters / 3.2) and marked `estimated` in the record.
4. Not implemented: "one object per write call" rule and splitting large test classes. Both change what the model sees (system prompt or task) and so the training distribution and
the comparison with the baseline (same system prompt). Option if 1 to 3 are not enough: a hint in the task spec of CDS tasks only.
**Verify after the reset (5 runs, DDLS and PROG first because they had most empty runs).** Run with `STREAM_GUARD=9000 python3 -m harness.pipeline` or set the constant:
check that `empty_response` stays below 5 % over 40 runs, that no accepted run was cut wrongly (`cut_by_stream_guard` in `metadata.turn_usage` followed by a good turn is fine),
and that the cost per run does not rise. If the stream breaks (tool calls missing), unset `STREAM_GUARD`: items 1 and 2 stay.

View File

@@ -1,4 +1,4 @@
# EPOD: stale lock after a write (2 cases, not reproduced under control)
# EPOD: stale lock after a write (3 cases; the third one gives the cause)
Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide that it is not worth it).
@@ -13,6 +13,24 @@ Status 2026-10-05 evening. For Kral, to fix in the EPOD server (or to decide tha
In both cases the lock outlived the MCP session and the run (hours). `ENQUEUE_READ` (see below) showed nothing after the SM12 delete.
## Case 3 and the probable cause (2026-10-06): the MCP server (Eclipse) restarted while a write was running
At 21:43 on 2026-10-05 Kral restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment. One of them (run 202781, task G1927) left this entry in the
enqueue table (read with the reader class below, 2026-10-06 06:15):
```
GNAME=SEOCLSENQ GARG=ZCL_Z4CGT1HJ_JOB_COST_TEST====... GMODE=X GOBJ=ESEOCLASS GCLIENT=001 GUNAME=KESELI
GUSR=20261005194600195642000500vhcala4hci_A4H_00... GUSE=1 GTHOST=vhcala4hci_A4H_00 GTWP=05 GTDATE=20261005 GTTIME=194600 (server time, 2 hours behind CEST: 21:46)
```
Exclusive lock (mode X) on the class, owner user KESELI, created at 21:46, three minutes after the restart, by a write that was in flight; it was still there the next morning.
Its object was not cleaned up because the model had named it `ZCL_<run prefix>_JOB_COST_TEST` (prefix inside the name; the teardown looked only for names that start with the prefix). When the pipeline
restarted at 22:27 it ran the same task with the same run number, so the new run found the old locked class: `[LOCK] ... User KESELI is currently editing` on every `sap_push_element` (the run looped and scored 75).
Earlier I wrote that the restart did not leak a lock: that was wrong, I had only cleaned the objects whose names start with the prefix.
So the cause is probably: **a write call is in flight when the MCP server process stops (restart, crash, kill of the whole server); the lock of that write stays in the enqueue table** (the stateful ADT session that owns it is gone, and nothing removes the lock).
This also fits case 2 (two controllers killed in the middle of a write) better than the client-side kill tests, which did not reproduce it: killing the *client* does not stop the server's write.
It cannot be tested from the client side without restarting Eclipse; for Kral: start a write (a class with a large test include), restart Eclipse in the middle, read the enqueue table (below) and look at SM12.
## What we tried to reproduce (A4H, probe objects, no cloud, 2026-10-05)
Every test: write through MCP, interfere, wait 3 s, read the enqueue table, write the same object again, delete it.
@@ -53,8 +71,13 @@ There is no function module on A4H to delete an enqueue entry (`TFDIR`: only `EN
- One controller only (flock on `runs/pipeline/controller.lock`), so two controllers cannot install the same run number twice any more (this was Case 2).
- Never kill a controller while it runs: it is drained through the budget guard or a STOP flag.
## Also found: leftovers with the prefix inside the name
27 classes named `ZCL_<prefix>_...` / `ZCX_<prefix>_...` were left in A4H from earlier runs (the teardown searched only for names that start with the prefix). They showed up in
`sap_inactive_objects` and searches of later runs and in 57 of 90 accepted trajectories. Fixed in the harness (2026-10-06, `harness/proxy.py`, `harness/runner.py`, `harness/sweep.py`).
## Wish for EPOD
0. **Release the enqueue locks of the server's own ADT sessions when the MCP server stops or starts** (a lock of user KESELI with the server's work process that has no live ADT session behind it).
1. Release the lock in a `finally` path of every write tool (also after a failed or warned activation).
2. Release all locks of an MCP session when the session ends (`DELETE /mcp`) and when the connection breaks.
3. Return a clear error with the lock owner and the age of the lock, and a tool to release the locks of the caller's own session.

View File

@@ -0,0 +1,44 @@
# Eval seti: yeni türler için slot planı ve Kral'ın kontrol sayfası
Durum: hazırlık (2026-10-06). Üretim 12 Ekim sıfırlamasından sonra, **eğitim görevlerinden önce** (`docs/restart-12-oktober.md`, adım 0).
Amaç: eğitimden sonra INTF, TABL, STRU, MSAG ve istisna sınıfı görevlerini ölçebilmek. Şu an eval adaylarında (130 kabul) sadece CLAS, FUNC, DDLS, PROG var.
## 1. Slotlar (`python3 -m harness.evalset plan-new`)
Eğitim karışımıyla aynı paylar (110 görevlik sette INTF 8, TABL 9, STRU 2, MSAG 3, istisna 3); %30 fazla aday, çünkü ret ve ampirik süzgeç göreve göre eler.
| tür | kategori | slot | kimlikler | not |
|---|---|---|---|---|
| INTF | A | 10 | G0200–G0209 | arayüz: sabitler, tipler, metotlar; gizli test arayüzü uygulayan yerel sınıf kullanır |
| TABL | B | 11 | G0210–G0220 | tablo: alanlar, anahtar, not null; gizli test RTTI + `cl_osql_test_environment` |
| STRU | B | 3 | G0221–G0223 | yapı: alanlar, tipler; RTTI |
| MSAG | D | 4 | G0224–G0227 | mesaj sınıfı: numara, metin, yer tutucu; `MESSAGE ... INTO` |
| istisna (CLAS) | D | 4 | G0228–G0231 | `CX_` sınıfı: öznitelik, metin, zincir |
| K (serbest metin) | K | 5 | G0232–G0236 | INTF 1, TABL 2, MSAG 1, istisna 1; `evalset run-new-k`, tool isimleri EPOD'daki gibi (`generic_v0` kullanılmıyor, varsayım: bunu Kral onaylar) |
Her ikinci-üçüncü slot zor (zorluk 3). Her slot eğitim havuzuyla örtüşme kontrolünden geçer (eğitimdeki görevlere çok benziyorsa model yeniden yazar).
Maliyet tahmini: 32 slot x yaklaşık 0,15 ledger = 5 ledger, yaklaşık 2 usage; K varyantları 1 ledger'dan az.
## 2. Otomatik kapılar (Claude, senin işin değil)
oracle 100 ve null 0, mutasyon denetimi (INTF/TABL/STRU/MSAG/istisna için yeni mutantlar), eğitim havuzuyla örtüşme, ampirik süzgeç (DeepSeek ve yerel Qwen: seri A düzeneği eval görevlerini de koşturur),
sonra `docs/eval-inceleme.md` kontrol listesiyle Claude incelemesi (`python3 -m harness.review G0200`).
## 3. Senin kontrolün (yaklaşık 10 görev, yaklaşık 45 dakika)
Kimler: Claude'un **işaretlediği** görevler + her yeni türden **1 örnek** (5 görev) + 1 K varyantı. Her görev için `python3 -m harness.review <id>` tek ekranda spec, sözleşme, gizli test isimleri, mutasyon sonucu ve üretim geçmişini gösterir.
Kararın: `python3 -m harness.review <id> --set accept|fix|flag|reject --note "..."`.
Genel sorular (listedeki 1–11): spec belirsiz mi, her kuralın testi var mı, kontrat yeterli mi, zanaat serbest mi, gerçekçi mi, sızıntı var mı.
**Yeni türlere özel sorular (12–16):**
| # | Soru | tür |
|---|---|---|
| 12 | Spec her alanı, uzunluğu, anahtarı ve not-null'ı tek anlamlı veriyor mu? Gizli test **bayt/karakter** karışıklığına düşmüyor mu (RTTI `length` bayttır)? | TABL, STRU |
| 13 | Tablo testi gerçek davranışı ölçüyor mu (aynı anahtarla ikinci INSERT `sy-subrc 4`), yoksa sadece alan adlarına mı bakıyor? | TABL |
| 14 | Arayüz testi imza uyuşmazlığını yakalıyor mu (uygulayan yerel sınıf yalnızca imza doğruysa derlenir) ve sabit değerlerini okuyor mu? | INTF |
| 15 | Mesaj metinleri, numaralar ve yer tutucular spec'te **birebir** yazılı mı; test `MESSAGE ID ... INTO` ile son metni mi karşılaştırıyor? | MSAG |
| 16 | İstisna: öznitelik, metin ve (varsa) önceki istisna zinciri spec'te tanımlı mı; test fırlat-yakala ile ölçüyor mu? Gerçekten "istisna tasarımı" mı gerekiyor? | istisna |
## 4. Sonuç tablosu (üretimden sonra doldurulur)
| id | tür | otomatik kapılar | Claude incelemesi | Kral | not |
|---|---|---|---|---|---|
| | | | | | |

View File

@@ -16,6 +16,13 @@ Ollama reset on 12 October. Until then only no-cloud work. This is the plan for
Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist.
4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`).
## 1b. Step 0: eval tasks for the new kinds (before any training task)
`python3 -m harness.evalset plan-new` shows 32 slots (INTF 10, TABL 11, STRU 3, MSAG 4, exception 4; ids G0200 to G0231), then `python3 -m harness.evalset run-new`
(about 5 ledger, about 2 usage), then `python3 -m harness.evalset run-new-k` (5 K variants). Reason: the eval set (130 accepted candidates) has no INTF, TABL, STRU, MSAG
or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: `docs/eval-spotcheck-new-kinds.md`
(Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.
## 2. Order of work (what the controller does by itself)
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest

58
docs/stage2-build.md Normal file
View File

@@ -0,0 +1,58 @@
# Stage 2 training set, build of 2026-10-06 06:15:28
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
## Result
81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
Dropped:
- FUNC: {'read_of_another_runs_object': 3}
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
- TABL: {'read_of_another_runs_object': 1}
- INTF: {'over_48k': 1}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
## Not good enough yet (kinds under the minimum of 25)
| kind | samples | families | short |
|---|---|---|---|
| CLAS | 14 | 14 | 11 |
| INTF | 2 | 2 | 23 |
| DDLS | 9 | 9 | 16 |
| FUNC | 8 | 8 | 17 |
| PROG | 5 | 5 | 20 |
| TABL | 2 | 2 | 23 |
| STRU | 0 | 0 | 25 |
| MSAG | 0 | 0 | 25 |
| EXC | 0 | 0 | 25 |
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides)
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
| 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M |
| 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M |
| 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M |
| 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M |
| 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M |
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).

22
epod_tests/README.md Normal file
View File

@@ -0,0 +1,22 @@
# EPOD acceptance tests (for the EPOD server developer)
Small test set for the two server changes that the harness work asked for (details: `docs/epod-syntax-hint.md`, `docs/epod-lock-leak.md`)
and for the parallel-call behavior (`harness/loadtest.py` numbers). Standard library only, Python 3.9+, no harness code needed.
It uses probe objects `ZEPODT_*` in `$TMP` on a **test system**.
```sh
cd epod_tests
export MCP_URL=http://127.0.0.1:3000/mcp MCP_TOKEN=... # the MCP server
export A4H_URL=http://localhost:50000 A4H_USER=... A4H_PASSWORD=... # only for the cleanup (ADT deletion); without it the tests list the probe objects
python3 run_tests.py # all; --only T1 T3 for some
```
| test | passes when | state of the server on 2026-10-06 |
|---|---|---|
| T1 syntax_hint | a rejected write (`TYPE c LENGTH 4` in a method signature) returns syntax messages with a line, not only "save operation failed" | FAIL expected (not implemented; the harness proxy adds abaplint messages) |
| T2 no_lock_after_kill | a client killed with SIGKILL during a write (4 delays) leaves no lock: the next write works | PASS (not reproducible on A4H) |
| T3 parallel_reads | 6 clients search + syntax check in parallel without `Concurrent call detected` | PASS |
| T4 parallel_writes | 3 clients create + write in parallel without `Concurrent call detected` | FAIL expected (the shared connection; 15 to 50 lock hits per 45 s in the harness load test) |
| T5 lock_error_text | informational | PASS |
A change in the server is done when T1 and T4 pass and T2 and T3 stay green. T1 accepts any answer that carries a line number or a `syntaxCheck` object with messages (the proxy format is in `docs/epod-syntax-hint.md`).

98
epod_tests/epod_client.py Normal file
View File

@@ -0,0 +1,98 @@
"""Minimal MCP client and ADT deletion for the EPOD acceptance tests (standard library only, Python 3.9+)."""
import base64
import http.cookiejar
import json
import os
import re
import time
import urllib.request
from xml.sax.saxutils import quoteattr
URL = os.environ.get("MCP_URL", "http://127.0.0.1:3000/mcp")
TOKEN = os.environ.get("MCP_TOKEN", "")
SYSTEM = os.environ.get("MCP_SYSTEM_ID") # optional: Eclipse project name when several systems are connected
class Mcp:
def __init__(self, timeout=300):
self.sid, self.n, self.timeout = None, 0, timeout
def _post(self, body, method="POST"):
h = {"Content-Type": "application/json", "Accept": "application/json, text/event-stream"}
if TOKEN:
h["Authorization"] = "Bearer " + TOKEN
if self.sid:
h["Mcp-Session-Id"] = self.sid
req = urllib.request.Request(URL, json.dumps(body).encode() if body is not None else None, h, method=method)
with urllib.request.urlopen(req, timeout=self.timeout) as r:
sid = r.headers.get("Mcp-Session-Id")
raw = r.read().decode()
self.sid = sid or self.sid
if "data:" in raw[:40]:
raw = "".join(l[5:].strip() for l in raw.splitlines() if l.startswith("data:"))
return json.loads(raw) if raw.strip() else None
def open(self):
self.n += 1
self._post({"jsonrpc": "2.0", "id": self.n, "method": "initialize", "params": {
"protocolVersion": "2025-03-26", "capabilities": {}, "clientInfo": {"name": "epod-acceptance", "version": "1"}}})
self._post({"jsonrpc": "2.0", "method": "notifications/initialized"})
return self
def close(self):
try:
self._post(None, "DELETE")
except Exception:
pass
def __enter__(self):
return self.open()
def __exit__(self, *a):
self.close()
def call(self, tool, args):
"""Returns (is_error, text). 'Concurrent call detected' is returned as it is (the tests count it)."""
self.n += 1
if SYSTEM:
args = dict(args, systemId=SYSTEM)
res = self._post({"jsonrpc": "2.0", "id": self.n, "method": "tools/call", "params": {"name": tool, "arguments": args}})
if "error" in res:
return True, json.dumps(res["error"])
r = res["result"]
return bool(r.get("isError")), "\n".join(c.get("text", "") for c in r.get("content", []))
class Adt:
"""Deletion through the ADT deletion API (needs A4H_URL, A4H_USER, A4H_PASSWORD; A4H_CLIENT default 001)."""
def __init__(self):
self.base = os.environ.get("A4H_URL", "").rstrip("/")
self.client = os.environ.get("A4H_CLIENT", "001")
self.auth = "Basic " + base64.b64encode(("%s:%s" % (os.environ.get("A4H_USER", ""), os.environ.get("A4H_PASSWORD", ""))).encode()).decode()
self.opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(http.cookiejar.CookieJar()))
self.csrf = None
def available(self):
return bool(self.base and os.environ.get("A4H_USER"))
def _req(self, path, body=None, headers=None):
h = {"Authorization": self.auth}
if self.csrf:
h["x-csrf-token"] = self.csrf
h.update(headers or {})
req = urllib.request.Request("%s%s%ssap-client=%s" % (self.base, path, "&" if "?" in path else "?", self.client), body.encode() if body else None, h)
with self.opener.open(req, timeout=120) as r:
return dict(r.headers), r.read().decode()
def delete(self, uris):
if not self.available() or not uris:
return {}
hdr, _ = self._req("/sap/bc/adt/discovery", headers={"x-csrf-token": "fetch", "Accept": "*/*"})
self.csrf = hdr.get("x-csrf-token") or hdr.get("X-CSRF-Token")
objs = "".join("<del:object adtcore:uri=%s><del:transportNumber/></del:object>" % quoteattr(u) for u in uris)
body = ('<?xml version="1.0" encoding="UTF-8"?><del:deletionRequest xmlns:del="http://www.sap.com/adt/deletion" '
'xmlns:adtcore="http://www.sap.com/adt/core">' + objs + "</del:deletionRequest>")
_, text = self._req("/sap/bc/adt/deletion/delete", body, {"Content-Type": "application/vnd.sap.adt.deletion.request.v1+xml",
"Accept": "application/vnd.sap.adt.deletion.response.v1+xml"})
return {m.group(1): 'isDeleted="true"' in m.group(0) for m in re.finditer(r'<del:object\b[^>]*adtcore:uri="([^"]+)"[^>]*>', text)}

View File

@@ -0,0 +1,32 @@
[
{
"id": "T1",
"name": "syntax_hint",
"pass": false,
"detail": "the failed write returned: {\"success\":false,\"error\":\"[WRITE] An error occured during the save operation. The changes were not stored.\"}"
},
{
"id": "T2",
"name": "no_lock_after_kill",
"pass": true,
"detail": "4 kills (0.05 to 0.7 s), the next write always worked"
},
{
"id": "T3",
"name": "parallel_reads",
"pass": true,
"detail": "372 rounds, 0 'Concurrent call detected'"
},
{
"id": "T4",
"name": "parallel_writes",
"pass": false,
"detail": "27 create+write rounds with 3 clients, 12 'Concurrent call detected'"
},
{
"id": "T5",
"name": "lock_error_text",
"pass": true,
"detail": "informational: see docs/epod-lock-leak.md (no way to create a lock on purpose)"
}
]

174
epod_tests/run_tests.py Normal file
View File

@@ -0,0 +1,174 @@
"""EPOD acceptance tests (2026-10-06). Run against an EPOD MCP server on a test system (probe objects ZEPODT_*, package $TMP).
MCP_URL=http://127.0.0.1:3000/mcp MCP_TOKEN=... [A4H_URL=http://localhost:50000 A4H_USER=... A4H_PASSWORD=...] python3 run_tests.py [--only T1 T2 ...]
Each test prints PASS or FAIL with the reason and writes the result to epod_tests_result.json. Tests:
T1 syntax_hint a rejected write ('save operation failed') comes with syntax messages (line and text) docs/epod-syntax-hint.md
T2 no_lock_after_kill a client killed during a write leaves no lock (the next write and the deletion work) docs/epod-lock-leak.md
T3 parallel_reads 6 clients read/ATC/syntax-check in parallel without 'Concurrent call detected'
T4 parallel_writes 3 clients create/write/activate in parallel without 'Concurrent call detected' (a fix: per system or per session lock)
T5 lock_error_text a write on a locked object names the lock owner (informational)
"""
import json
import os
import signal
import subprocess
import sys
import threading
import time
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from epod_client import Adt, Mcp # noqa: E402
CLS = """CLASS {n} DEFINITION PUBLIC FINAL CREATE PUBLIC.
PUBLIC SECTION.
METHODS run RETURNING VALUE(rv) TYPE i.
ENDCLASS.
CLASS {n} IMPLEMENTATION.
METHOD run.
rv = 1.
ENDMETHOD.
ENDCLASS.
"""
BAD = """CLASS {n} DEFINITION PUBLIC FINAL CREATE PUBLIC.
PUBLIC SECTION.
METHODS run IMPORTING iv_zone TYPE c LENGTH 4 RETURNING VALUE(rv) TYPE i.
ENDCLASS.
CLASS {n} IMPLEMENTATION.
METHOD run.
rv = 1.
ENDMETHOD.
ENDCLASS.
"""
TC = """CLASS ltc DEFINITION FINAL FOR TESTING DURATION SHORT RISK LEVEL HARMLESS.
PRIVATE SECTION.
METHODS t1 FOR TESTING.
ENDCLASS.
CLASS ltc IMPLEMENTATION.
METHOD t1.
cl_abap_unit_assert=>assert_equals( act = 1 exp = 1 ).
ENDMETHOD.
ENDCLASS.
"""
created = []
RUN = time.strftime("%H%M%S")
def new_class(m, tag):
name = "ZEPODT_%s_%s" % (tag, RUN)
m.call("sap_create_object", {"objectType": "CLAS", "objectName": name, "packageName": "$TMP", "description": "epod acceptance test"})
created.append("/sap/bc/adt/oo/classes/" + name.lower())
return name
def ok(text):
return '"success":true' in text.replace(" ", "")
def t1_syntax_hint():
with Mcp() as m:
n = new_class(m, "T1")
e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": n, "source": BAD.format(n=n.lower())})
low = t.lower()
has_detail = ("syntaxcheck" in low or "syntax" in low and "line" in low) and "save operation" in low or "line 3" in low or '"line":3' in low.replace(" ", "")
return bool(has_detail), "the failed write returned: " + t[:300]
def t2_no_lock_after_kill():
victim = ("import sys,json,os\nsys.path.insert(0,%r)\nfrom epod_client import Mcp\na=json.loads(sys.argv[1])\nm=Mcp().open()\nprint('SENT',flush=True)\nm.call('sap_push_source',a)\n"
% os.path.dirname(os.path.abspath(__file__)))
bad = []
for delay in (0.05, 0.2, 0.4, 0.7):
with Mcp() as m:
n = new_class(m, "T2")
m.call("sap_push_source", {"objectType": "CLAS", "objectName": n, "source": CLS.format(n=n.lower())})
p = subprocess.Popen([sys.executable, "-c", victim, json.dumps({"objectType": "CLAS", "objectName": n, "includeType": "testclasses", "source": TC})],
stdout=subprocess.PIPE, text=True)
p.stdout.readline()
time.sleep(delay)
try:
os.kill(p.pid, signal.SIGKILL)
except ProcessLookupError:
pass
p.wait()
time.sleep(3)
with Mcp() as m:
e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": n, "includeType": "testclasses", "source": TC + "* again\n"})
if not ok(t):
bad.append((delay, t[:160]))
return not bad, "after the kill the next write failed at: %s" % bad if bad else "4 kills (0.05 to 0.7 s), the next write always worked"
def parallel(n_clients, seconds, work):
out, end = [], time.time() + seconds
def run(i):
locks = calls = 0
with Mcp() as m:
while time.time() < end:
t = work(m, i)
calls += 1
locks += "Concurrent call detected" in t
out.append((calls, locks))
th = [threading.Thread(target=run, args=(i,)) for i in range(n_clients)]
[x.start() for x in th]
[x.join() for x in th]
return sum(c for c, _ in out), sum(l for _, l in out)
def t3_parallel_reads():
def work(m, i):
return m.call("sap_search_object", {"query": "CL_ABAP_CHAR_UTIL*", "objType": "CLAS"})[1] + \
m.call("sap_syntax_check", {"objectType": "CLAS", "objectName": "CL_ABAP_CHAR_UTILITIES"})[1]
calls, locks = parallel(6, 15, work)
return locks == 0, "%d rounds, %d 'Concurrent call detected'" % (calls, locks)
def t4_parallel_writes():
cnt = {"n": 0}
def work(m, i):
cnt["n"] += 1
name = "ZEPODT_P%d_%s%03d" % (i, RUN, cnt["n"] % 1000)
t = m.call("sap_create_object", {"objectType": "CLAS", "objectName": name, "packageName": "$TMP", "description": "parallel"})[1]
created.append("/sap/bc/adt/oo/classes/" + name.lower())
return t + m.call("sap_push_source", {"objectType": "CLAS", "objectName": name, "source": CLS.format(n=name.lower())})[1]
calls, locks = parallel(3, 25, work)
return locks == 0, "%d create+write rounds with 3 clients, %d 'Concurrent call detected'" % (calls, locks)
def t5_lock_error_text():
return True, "informational: see docs/epod-lock-leak.md (no way to create a lock on purpose)"
TESTS = [("T1", "syntax_hint", t1_syntax_hint), ("T2", "no_lock_after_kill", t2_no_lock_after_kill), ("T3", "parallel_reads", t3_parallel_reads),
("T4", "parallel_writes", t4_parallel_writes), ("T5", "lock_error_text", t5_lock_error_text)]
def main():
only = set(sys.argv[sys.argv.index("--only") + 1:]) if "--only" in sys.argv else None
result = []
for tid, name, fn in TESTS:
if only and tid not in only:
continue
t0 = time.time()
try:
passed, why = fn()
except Exception as e: # noqa: BLE001
passed, why = False, "error: %r" % e
print("%-4s %-22s %s (%.0f s) %s" % (tid, name, "PASS" if passed else "FAIL", time.time() - t0, why[:230]), flush=True)
result.append({"id": tid, "name": name, "pass": passed, "detail": why[:600]})
adt = Adt()
if adt.available() and created:
try:
gone = adt.delete(created)
print("cleanup: %d of %d probe objects deleted" % (sum(gone.values()), len(created)))
except Exception as e: # noqa: BLE001
print("cleanup failed:", repr(e)[:200])
elif created:
print("delete these probe objects yourself (SE80 or ADT): ZEPODT_* in $TMP (%d objects)" % len(created))
json.dump(result, open("epod_tests_result.json", "w"), indent=1)
if __name__ == "__main__":
main()

View File

@@ -84,7 +84,7 @@ class LlmAgent:
def __init__(self, model, base_url=None, api_key=None, max_turns=80, temperature=0.2,
max_seconds=None, max_tokens=None, chat_template_kwargs=None, loop_guard=None,
deadline=None, watch=False):
deadline=None, watch=False, empty_retries=2, retry_temperature=None, stream_guard=None):
self.model = model
self.name = f"llm:{model}"
self.base_url = (base_url or os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1")).rstrip("/")
@@ -96,7 +96,12 @@ class LlmAgent:
# Runaway reasoning: 35 of 1965 DeepSeek turns produced 393k output tokens and no content (46 % of the
# run cost, 2026-10-03). Normal turns: p95 17.5k, max 82k. A cut turn is retried (see run).
self.max_tokens = max_tokens or (32000 if ":cloud" in (model or "") else None)
self.empty_retries = 2
self.empty_retries = empty_retries
# empty_response fix (2026-10-06): a retry after an empty turn can use another temperature (the same sample often
# runs away again: 7 of 13 empty runs had two or three capped turns in a row); stream_guard = reasoning tokens
# after which a streamed turn without any content or tool call is cut and counted as an empty turn
self.retry_temperature = retry_temperature
self.stream_guard = stream_guard
self.loop_guard = loop_guard # end the run after this many identical pushes in a row (None = off)
self.deadline = deadline # absolute time (time.time()) of the hard stop, or None
self.watch = watch # remote model: ping the server during a request, end the run when it is gone
@@ -104,9 +109,9 @@ class LlmAgent:
self.messages, self.tools, self.reasoning, self.turn_usage = [], [], [], [] # for the trajectory record
self.chat_template_kwargs = chat_template_kwargs # local server only, e.g. {"enable_thinking": False}
def _chat(self, messages, tools):
def _chat(self, messages, tools, temperature=None):
body = {"model": self.model, "messages": messages, "tools": tools,
"temperature": self.temperature, "parallel_tool_calls": False}
"temperature": temperature if temperature is not None else self.temperature, "parallel_tool_calls": False}
if self.max_tokens:
body["max_tokens"] = self.max_tokens
if self.chat_template_kwargs:
@@ -115,6 +120,11 @@ class LlmAgent:
{"Content-Type": "application/json",
"Authorization": f"Bearer {self.api_key}"})
last = None
if self.stream_guard and not (self.watch or self.deadline):
try:
return self._chat_stream(dict(body), messages)
except Exception as e: # noqa: BLE001 streaming not usable (server, parse): the normal request below
last = e
for attempt in range(4): # model server errors (HTTP 5xx, timeouts): retry with backoff
try:
if self.watch or self.deadline:
@@ -137,6 +147,49 @@ class LlmAgent:
time.sleep(10 * (attempt + 1))
raise RuntimeError(f"model request failed after retries: {last}")
def _chat_stream(self, body, messages):
"""Streamed request. A turn that has produced only reasoning for `stream_guard` tokens (about 3.2 characters per
token) and no content and no tool call is cut and returned as an empty turn (the run loop retries it)."""
body["stream"] = True
body["stream_options"] = {"include_usage": True}
req = urllib.request.Request(f"{self.base_url}/chat/completions", json.dumps(body).encode(),
{"Content-Type": "application/json", "Authorization": f"Bearer {self.api_key}"})
content, reasoning, calls, usage, cut = "", 0, {}, {}, False
with urllib.request.urlopen(req, timeout=self.request_timeout) as r:
for raw in r:
line = raw.decode("utf-8", "replace").strip()
if not line.startswith("data:"):
continue
data = line[5:].strip()
if data == "[DONE]":
break
chunk = json.loads(data)
if chunk.get("usage"):
usage = chunk["usage"]
for ch in chunk.get("choices") or []:
d = ch.get("delta") or {}
content += d.get("content") or ""
reasoning += len(d.get("reasoning") or d.get("reasoning_content") or d.get("thinking") or "")
for tc in d.get("tool_calls") or []:
c = calls.setdefault(tc.get("index", 0), {"id": tc.get("id") or "", "type": "function",
"function": {"name": "", "arguments": ""}})
c["id"] = c["id"] or tc.get("id") or ""
f = tc.get("function") or {}
c["function"]["name"] += f.get("name") or ""
a = f.get("arguments")
c["function"]["arguments"] += a if isinstance(a, str) else json.dumps(a) if a else ""
if not content.strip() and not calls and reasoning / 3.2 >= self.stream_guard:
cut = True
break
if cut:
est = {"prompt_tokens": int(len(json.dumps(messages)) / 3.5), "completion_tokens": int(reasoning / 3.2),
"estimated": True, "cut_by_stream_guard": True}
return {"role": "assistant", "content": ""}, est
msg = {"role": "assistant", "content": content or None}
if calls:
msg["tool_calls"] = [calls[k] for k in sorted(calls)]
return msg, usage
def _ping(self, timeout=10):
try:
urllib.request.urlopen(urllib.request.Request(f"{self.base_url}/models"), timeout=timeout).read()
@@ -195,7 +248,8 @@ class LlmAgent:
self.end_reason = "time_budget"
break
try:
msg, usage = self._chat(messages, tools)
msg, usage = self._chat(messages, tools,
self.retry_temperature if (empty and self.retry_temperature) else None)
add_usage(self.model, usage, kind="run", ref=proxy.prefix)
self.turn_usage.append(usage)
except WindowEnd:

View File

@@ -302,6 +302,22 @@ def local_card():
return hdr + tbl
def work_card():
"""Progress of the no-cloud work packages (runs/dashboard/work.json, updated by Claude after each item)."""
path = os.path.join(OUT_DIR, "work.json")
if not os.path.exists(path):
return ""
try:
w = json.load(open(path))
except ValueError:
return ""
cls = {"done": "ok", "in progress": "wa", "waiting": "gr", "parked": "gr", "blocked": "er"}
rows = "".join("<tr><td><b>%s</b></td><td>%s</td><td><span class='chip %s'>%s</span></td><td>%s</td></tr>" % (
E(i["id"]), E(i["text"]), cls.get(i["status"], "gr"), E(i["status"]), E(i.get("note", ""))) for i in w.get("items", []))
return ("<div class=card style='grid-column:1/-1'><h2>%s</h2><div class=b><table><tr><th>item</th><th>work</th><th>status</th><th>result</th></tr>%s</table></div>"
"<div class=note>updated %s</div></div>" % (E(w.get("title", "")), rows, time.strftime("%H:%M", time.localtime(w.get("updated", 0)))))
def mix_table(rows, evs):
kt = mix.accepted_task_counts()
kr = {}
@@ -437,7 +453,7 @@ def render(d):
% ("ok" if d["a4h"] else "er", "up" if d["a4h"] else "down", "ok" if d["mcp"] else "er", "up" if d["mcp"] else "down", scls, status,
E("\n".join(d["procs"]))))
logc = "<div class=card><h2>Pipeline log</h2><div class=b><div class='log mono'>%s</div></div></div>" % E("\n".join(d["log_tail"]))
body = (banner + "<div class=tiles>" + tiles + "</div><div class=grid>" + local_card() + budget + prog + stops + sysc
body = (banner + "<div class=tiles>" + tiles + "</div><div class=grid>" + work_card() + local_card() + budget + prog + stops + sysc
+ mix_table(rows, evs) + agg_table("Trajectories by category", evs, rows, "category") + agg_table("Trajectories by object type", evs, rows, "object_type")
+ gen_table(logs, "category", "Task generation by category") + gen_table(logs, "object_type", "Task generation by object type")
+ tok + recent + logc + "</div>")

View File

@@ -16,6 +16,7 @@ import sys
from .adt_client import load_env
from .generator import ROOT, generate, make_k_variant
from . import mix, overlap
from .ledger import BudgetExceeded, spent
FIRST_ID = 100
@@ -52,6 +53,73 @@ K_SOURCES = [("G0004", "free_text"), ("G0007", "incomplete"), ("G0026", "free_te
RELEASES = ["v702", "v740sp05"]
# New kinds for the eval set (Opus item F, 2026-10-06): the 130 eval candidates have only CLAS, FUNC, DDLS and PROG, so INTF, TABL, STRU, MSAG and
# exception tasks cannot be measured after training. Shares as in the training mix (110 tasks: INTF 8, TABL 9, STRU 2, MSAG 3, exception 3); about
# 30 percent more slots than the target because of rejections and the empirical filter. Generation after the reset, FIRST (before training tasks).
NEW_FIRST_ID = 200
NEW_RUN_BASE = 440000 # 40 per slot, below 466560
NEW_KIND_SLOTS = [("INTF", "A", 10), ("TABL", "B", 11), ("STRU", "B", 3), ("MSAG", "D", 4), ("EXC", "D", 4)]
NEW_K = [("INTF", "free_text"), ("TABL", "incomplete"), ("TABL", "free_text"), ("MSAG", "free_text"), ("EXC", "incomplete")]
EXC_TOPICS = ["exception class (CX_...): a domain exception with context attributes and message texts",
"exception class (CX_...): an exception hierarchy with a common super class",
"exception class (CX_...): an exception that wraps a previous exception",
"exception class (CX_...): an exception with a message class and parameters in the text"]
def plan_new():
out, n = [], 0
for kind, cat, count in NEW_KIND_SLOTS:
for i in range(count):
out.append({"id": f"G{NEW_FIRST_ID + n:04d}", "kind": kind, "category": cat, "object_type": "CLAS" if kind == "EXC" else kind,
"topic": EXC_TOPICS[i % len(EXC_TOPICS)] if kind == "EXC" else None,
"difficulty": 3 if i % 3 == 2 else 2, "run_base": NEW_RUN_BASE + 40 * n})
n += 1
return out
def run_new(only, model, base_url):
"""Generate the new-kind eval candidates. The bundle is also checked against the training pool (no near duplicate of a training task)."""
train = overlap.load_pool("train")
for s in plan_new():
if only and s["id"] not in only:
continue
if os.path.exists(os.path.join(ROOT, "tasks_gen", "eval", s["id"], "generation.json")):
continue
avoid = [g for g in accepted_goals() if g][-170:]
topic = ((s["topic"] + ". ") if s["topic"] else "Choose a new, realistic business topic. ") + "Do not repeat these existing topics: " + "; ".join(avoid)
def extra(b, _t=train):
return [f"Too close to task {i} (spec {sc['spec']:.2f}, rules {sc['core']:.2f}, names {sc['name']:.2f}). Choose another topic and other object names."
for i, sc in overlap.check(overlap.load_bundle(b), _t)[:3]]
try:
log = generate(s["id"], "eval", s["object_type"], s["category"], s["difficulty"], model, base_url, s["run_base"], topic, extra_check=extra)
except BudgetExceeded as e:
print("BUDGET", e, flush=True)
break
except Exception as e: # noqa: BLE001
log = {"id": s["id"], "error": str(e)[:500]}
log.update(kind=s["kind"], spent_total=spent())
print(json.dumps(log), flush=True)
def run_new_k(model, base_url):
"""K variants (free text or incomplete spec, EPOD tool names) of the first accepted eval task of each new kind."""
plan = plan_new()
first_k = NEW_FIRST_ID + len(plan)
for j, (kind, style) in enumerate(NEW_K):
src = next((s["id"] for s in plan if s["kind"] == kind and os.path.exists(os.path.join(ROOT, "tasks_gen", "eval", s["id"], "generation.json"))
and json.load(open(os.path.join(ROOT, "tasks_gen", "eval", s["id"], "generation.json"))).get("accepted")), None)
new_id = f"G{first_k + j:04d}"
if not src or os.path.exists(os.path.join(ROOT, "tasks_gen", "eval", new_id, "generation.json")):
continue
try:
log = make_k_variant(src, new_id, style, model, base_url, NEW_RUN_BASE + 40 * (len(plan) + j), tool_schema=None, pool="eval")
except BudgetExceeded as e:
print("BUDGET", e, flush=True)
break
print(json.dumps(log), flush=True)
def accepted_goals():
"""Goal lines of all generated tasks (eval and train), so the model does not repeat a topic."""
out = []
@@ -90,6 +158,18 @@ def plan():
def main():
load_env(os.path.join(ROOT, ".env"))
cmd = sys.argv[1] if len(sys.argv) > 1 else "plan"
model, base_url = "deepseek-v4.1-flash:cloud", os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1")
if cmd == "plan-new":
for x in plan_new():
print(x)
print(len(plan_new()), "slots,", len(NEW_K), "K variants after them")
return
if cmd == "run-new":
run_new(set(sys.argv[2:]), model, base_url)
return
if cmd == "run-new-k":
run_new_k(model, base_url)
return
slots = plan()
if cmd == "plan":
for s in slots:

200
harness/owntests.py Normal file
View File

@@ -0,0 +1,200 @@
"""Own-test mutation score of an accepted trajectory (Opus item D, 2026-10-06): metadata only, the acceptance filter does not change.
The model's own unit tests (a testclasses include of the contract class, or global test classes) are run against the faulty references
of the task (`faulty/`: mutants of the reference that the hidden tests kill). Per trajectory: one run with the correct reference (the tests must
pass there, otherwise they encode model specific behavior and a kill proves nothing), then one run per mutant. Status per mutant:
killed (an own test fails), survived (all own tests pass), invalid (the mutant or the tests do not activate, or no own test ran).
score = killed / (killed + survived) over the valid mutants. PROG tasks are not supported (the tests live inside the program source).
python3 -m harness.owntests [--tasks G1000 ...] [--limit N] [--workers 1] writes runs/traj/<run>/own_test_mutation.json
"""
import argparse
import glob
import json
import os
import re
import shutil
import sys
import time
from .adt_client import load_env
from .agents import OracleAgent
from . import mix
from .runner import Runner
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
POOL = os.path.join(ROOT, "tasks_gen", "train")
WORK = os.path.join(ROOT, "runs", "owntests")
RUN_BASE = 480000 # + sequence number; below 466560 is NOT needed here: the run numbers stay under 466560 with 4-char base36
RUN_BASE = 420000
MAX_MUTANTS = 5
def calls_of(messages):
res = {m.get("tool_call_id"): m["content"] for m in messages if m["role"] == "tool"}
out = []
for m in messages:
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
out.append((c["function"]["name"], a, res.get(c.get("id"), "")))
return out
def model_tests(rec, task_meta):
"""{'include': {OBJECT: source}, 'global': {NAME: source}} of the model's last successful pushes."""
contract = {c["name"].upper() for c in task_meta.get("contract", [])}
last = {}
for tool, a, res in calls_of(rec["messages"]):
if tool == "sap_push_source" and a.get("source") and '"success":true' in (res or "").replace(" ", ""):
last[(str(a.get("objectName", "")).upper(), str(a.get("includeType") or "").lower(), a.get("objectType"))] = a["source"]
pre = rec["prefix"].upper()
seed_hidden = {o["name"].replace("{{P}}", pre).upper() for k in ("seed", "hidden_tests") for o in task_meta.get(k, [])}
out = {"include": {}, "global": {}}
for (name, inc, otype), src in last.items():
contract_names = {c.replace("{{P}}", pre).upper() for c in contract}
if inc == "testclasses" and name in contract_names:
out["include"][name] = src
elif otype == "CLAS" and not inc and name not in contract_names and name not in seed_hidden \
and re.search(r"FOR\s+TESTING", src, re.I) and re.search(r"^\s*CLASS\s+\S+\s+DEFINITION[^.]*FOR\s+TESTING", src, re.I | re.M):
out["global"][name] = src
return out
def placeholder(src, prefix):
return re.sub(re.escape(prefix), "{{P}}", re.sub(re.escape(prefix.lower()), "{{p}}", src), flags=re.I) if False else \
src.replace(prefix.upper(), "{{P}}").replace(prefix.lower(), "{{p}}")
def mutant_files(task_id):
"""[(reference object name, path)] of faulty/ (a K variant has none: the base task's)."""
meta = json.load(open(os.path.join(POOL, task_id, "task.json")))
d = os.path.join(POOL, meta.get("base_task") or task_id, "faulty")
return sorted(glob.glob(os.path.join(d, "m*_*")))
def derive(task_id, run_label, tests, prefix, mutant_path=None):
"""Build a temporary task: the reference (optionally with one mutated object) plus the model's own tests only."""
src_dir = os.path.join(POOL, task_id)
pool = os.path.join(WORK, "pool")
dst = os.path.join(pool, task_id)
shutil.rmtree(dst, ignore_errors=True)
shutil.copytree(src_dir, dst, ignore=shutil.ignore_patterns("faulty", "generation.json", "review.json", "mutation.json", "empirical.json"))
meta = json.load(open(os.path.join(dst, "task.json")))
pre = prefix.upper()
mutated = None
if mutant_path:
base = re.sub(r"^m\d+_", "", os.path.basename(mutant_path))
refs = []
for o in meta["reference"]:
f = o.get("file")
is_test = o["type"] == "CLAS" and f and re.search(r"^\s*CLASS\s+\S+\s+DEFINITION[^.]*FOR\s+TESTING",
open(os.path.join(dst, f)).read(), re.I | re.M)
if is_test:
continue # the reference's own global test class: not the model's
o = dict(o)
o.pop("testclasses_file", None) # the reference's local tests are out; the model's go in
if mutant_path and f and os.path.basename(f) == base:
shutil.copy(mutant_path, os.path.join(dst, f))
mutated = o["name"]
name_up = o["name"].replace("{{P}}", pre).upper()
if name_up in tests["include"]:
rel = "reference/_own_include_%s.abap" % re.sub(r"\W", "_", name_up)
open(os.path.join(dst, rel), "w").write(placeholder(tests["include"][name_up], prefix))
o["testclasses_file"] = rel
refs.append(o)
for i, (name, src) in enumerate(tests["global"].items()):
rel = "reference/_own_global_%d.clas.abap" % i
open(os.path.join(dst, rel), "w").write(placeholder(src, prefix))
refs.append({"type": "CLAS", "name": placeholder(name, prefix), "file": rel, "description": "own test class"})
meta["reference"] = refs
meta["budget"] = dict(meta.get("budget", {}), max_tool_calls=200)
json.dump(meta, open(os.path.join(dst, "task.json"), "w"), indent=1)
return pool, mutated
def run_one(pool, task_id, run_no):
runner = Runner(pool, os.path.join(WORK, "runs"))
rep, _ = runner.run(task_id, OracleAgent(), run_no, teardown=True)
own = rep.get("own_tests") or {}
g = rep.get("gates") or {}
return {"tests": own.get("tests", 0), "failures": own.get("failures", 0), "active": bool(g.get("G1_active")), "run": run_no}
def score_run(run_dir, seq):
rec = json.load(open(os.path.join(run_dir, "record.json")))
tid = rec["task"]["id"]
meta = json.load(open(os.path.join(POOL, tid, "task.json")))
out = {"task": tid, "run": rec["run"], "kind": mix.kind_of_task_dir(tid), "time": time.strftime("%F %T")}
if any(c.get("type") == "PROG" for c in meta.get("contract", [])):
return dict(out, status="not_supported", reason="PROG: the tests are inside the program source")
tests = model_tests(rec, meta)
if not tests["include"] and not tests["global"]:
return dict(out, status="no_own_tests", score=None, mutants=[])
muts = mutant_files(tid)[:MAX_MUTANTS]
if not muts:
return dict(out, status="no_mutants", score=None, mutants=[])
pre = rec["prefix"]
pool, _ = derive(tid, "base", tests, pre)
base = run_one(pool, tid, RUN_BASE + seq * 10)
out["reference_run"] = base
out["tests_pass_on_reference"] = base["active"] and base["tests"] > 0 and base["failures"] == 0
res = []
for k, mp in enumerate(muts):
pool, mutated = derive(tid, "m%d" % k, tests, pre, mp)
r = run_one(pool, tid, RUN_BASE + seq * 10 + 1 + k)
status = "invalid" if (not r["active"] or r["tests"] == 0) else ("killed" if r["failures"] > 0 else "survived")
res.append({"mutant": os.path.basename(mp), "object": mutated, "status": status, "tests": r["tests"], "failures": r["failures"]})
valid = [x for x in res if x["status"] != "invalid"]
killed = [x for x in res if x["status"] == "killed"]
out.update(status="scored", mutants=res, valid=len(valid), killed=len(killed),
score=round(len(killed) / len(valid), 2) if valid else None)
shutil.rmtree(os.path.join(WORK, "pool", tid), ignore_errors=True)
return out
def main():
load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser()
ap.add_argument("--tasks", nargs="*")
ap.add_argument("--limit", type=int)
ap.add_argument("--redo", action="store_true")
a = ap.parse_args()
os.makedirs(WORK, exist_ok=True)
rows = [json.loads(l) for l in open(os.path.join(ROOT, "runs", "traj", "summary.jsonl"))]
todo = []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-")
if not os.path.exists(os.path.join(p, "record.json")):
continue
if a.tasks and r["task"] not in a.tasks:
continue
if os.path.exists(os.path.join(p, "own_test_mutation.json")) and not a.redo:
continue
rec = json.load(open(os.path.join(p, "record.json")))
if acc.judge(rec, r, 80)[0]:
todo.append((r, p))
if a.limit:
todo = todo[:a.limit]
print(len(todo), "accepted trajectories to score", flush=True)
for i, (r, p) in enumerate(todo):
t0 = time.time()
try:
res = score_run(p, i)
except Exception as e: # noqa: BLE001
res = {"task": r["task"], "status": "error", "error": repr(e)[:300]}
json.dump(res, open(os.path.join(p, "own_test_mutation.json"), "w"), indent=1)
print(r["task"], res.get("status"), res.get("score"), "valid", res.get("valid"), "killed", res.get("killed"),
"ref_ok", res.get("tests_pass_on_reference"), "%.0fs" % (time.time() - t0), flush=True)
if __name__ == "__main__":
main()

View File

@@ -51,6 +51,9 @@ def activation_messages(text, is_error=False):
RUN_PREFIX = re.compile(r"^Z\d[0-9A-Z]{5,6}_", re.I)
# The model sometimes names its own helper or test class with the prefix inside the name (ZCL_Z4CGT1HJ_JOB_COST_TEST). Such objects of
# other runs stay behind in A4H and show up in lists and searches (57 of 90 accepted trajectories saw them, 2026-10-06).
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
class BudgetExceeded(Exception):
@@ -111,7 +114,10 @@ class ToolProxy:
def _foreign(self, name):
n = (name or "").upper()
return bool(RUN_PREFIX.match(n)) and not n.startswith(self.prefix)
if RUN_PREFIX.match(n):
return not n.startswith(self.prefix)
m = MID_PREFIX.match(n)
return bool(m) and m.group(1).upper() != self.prefix
def _filter(self, tool, text):
if tool == "sap_short_dumps": # only dumps of this run: other runs (and mutants) also write dumps
@@ -124,7 +130,7 @@ class ToolProxy:
data["count"] = len(data["dumps"])
return json.dumps(data)
return text
if tool not in ("sap_search_object", "sap_usage_references"):
if tool not in ("sap_search_object", "sap_usage_references", "sap_inactive_objects"):
return text
try:
data = json.loads(text)
@@ -148,7 +154,9 @@ class ToolProxy:
else:
self.calls += 1
name = str(args.get("objectName", "")).upper()
if tool in WRITE_TOOLS and self._foreign(name):
if (tool in WRITE_TOOLS or tool in ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info",
"sap_run_unit_test", "sap_check_object", "sap_syntax_check", "sap_atc_run")) \
and self._foreign(name):
result = (True, f"{name} is not available.")
else:
if tool in WRITE_TOOLS and (tool == "sap_activate" or args.get("activate", True)):

View File

@@ -153,9 +153,16 @@ class Runner:
return out
def _objects_with_prefix(self, mcp, prefix):
_, text = mcp.call("sap_search_object", {"query": prefix + "*", "maxResults": 200})
return [d for d in (_json(text) or []) if d.get("name", "").upper().startswith(prefix)
and d.get("objectType")] # skips STOB entries of CDS entities
"""Objects of the run: the name starts with the prefix, or contains it after a short type part (the model named its own
helper or test class ZCL_<prefix>_..., 27 such objects were left behind before 2026-10-06)."""
out = {}
for q in (prefix + "*", "*_" + prefix + "*"):
_, text = mcp.call("sap_search_object", {"query": q, "maxResults": 200})
for d in (_json(text) or []):
n = d.get("name", "").upper()
if d.get("objectType") and (n.startswith(prefix) or re.match(r"^[A-Z]{1,5}_" + re.escape(prefix), n)):
out[n] = d # skips STOB entries of CDS entities (no objectType)
return list(out.values())
def _source(self, mcp, otype, name, fg=None):
err, text = mcp.call("sap_pull_source", _obj_args(otype, name, fg))

89
harness/sweep.py Normal file
View File

@@ -0,0 +1,89 @@
"""Find (and delete) objects that harness runs left in A4H: names with a run prefix at the start (Z4CGT1HJ_X) or inside (ZCL_Z4CGT1HJ_X).
python3 -m harness.sweep dry run: list them by prefix
python3 -m harness.sweep --delete delete them (ADT deletion API, dependency order); a locked object is reported, not forced
python3 -m harness.sweep --keep-recent 60 do not touch prefixes of runs whose directory changed in the last 60 minutes (default 60)
Never run it while a controller or a series is writing: it can only judge by directory age.
"""
import argparse
import glob
import json
import os
import re
import time
from .adt_client import load_env
from .mcp_client import McpClient
from .runner import DELETE_ORDER, delete_uris
from .task import prefix_for
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
START = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
MID = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
def recent_prefixes(minutes):
keep = set()
for d in glob.glob(os.path.join(ROOT, "runs", "**", "*_G*_*"), recursive=True) + glob.glob(os.path.join(ROOT, "runs", "**", "*_T*_*"), recursive=True):
b = os.path.basename(d)
m = re.match(r"^(\d+)_([GT]\d+)_", b)
if m and time.time() - os.path.getmtime(d) < minutes * 60:
keep.add(prefix_for(int(m.group(1)), m.group(2)).upper())
return keep
def main():
load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser()
ap.add_argument("--delete", action="store_true")
ap.add_argument("--keep-recent", type=int, default=60)
a = ap.parse_args()
found = {}
with McpClient() as m:
# run prefixes are Z + a digit + 6 characters: one query per digit (a single "Z*" query is cut at 1000 results), at the start and after the type part
for q in [f"Z{d}*" for d in "0123456789"] + [f"*_Z{d}*" for d in "0123456789"] + ["ZPROBE*", "ZTEST*"]:
try:
for o in json.loads(m.call("sap_search_object", {"query": q, "maxResults": 1000})[1]):
if o.get("objectType"): # a STOB entry of a CDS view has the same name and no objectType: skip it
found[(o["name"], o["objectType"])] = o
except ValueError:
pass
keep = recent_prefixes(a.keep_recent)
by_prefix = {}
for (n, _t), o in found.items():
mm = START.match(n.upper()) or MID.match(n.upper())
if not mm or not o.get("objectType"):
continue
pre = mm.group(1).upper()
if len(pre) != 9 or not pre[1].isdigit():
continue
if pre in keep:
continue
by_prefix.setdefault(pre, []).append(o)
total = sum(len(v) for v in by_prefix.values())
print(f"{total} objects of {len(by_prefix)} run prefixes ({len(keep)} recent prefixes kept)")
for pre, objs in sorted(by_prefix.items()):
print(" ", pre, [o["name"] for o in objs][:6])
json.dump({p: [o["name"] for o in v] for p, v in by_prefix.items()}, open(os.path.join(ROOT, "runs", "sweep.json"), "w"), indent=1)
if not a.delete:
return
objs = sorted([o for v in by_prefix.values() for o in v],
key=lambda o: DELETE_ORDER.index(o["objectType"]) if o["objectType"] in DELETE_ORDER else 99)
done = failed = 0
for i in range(0, len(objs), 10):
uris = []
for o in objs[i:i + 10]:
uris.append(o["uri"])
res = delete_uris(uris)
for u, v in res.items():
if v["deleted"]:
done += 1
else:
failed += 1
print("not deleted:", u.split("/")[-1], v["msg"])
print(f"deleted {done}, not deleted {failed}")
if __name__ == "__main__":
main()

View File

@@ -28,7 +28,20 @@ POOL = os.path.join(ROOT, "tasks_gen", "train")
OUT = os.path.join(ROOT, "runs", "traj")
RUN_BASE = 200000 # 200000 + (task number - 1000) * 3 + attempt (a digit must lead the 4-char base36 run: < 466560)
MODEL = "deepseek-v4.1-flash:cloud"
CDS_CALLS = 100 # tool-call budget for tasks with a CDS contract object (eval keeps 60; Kral 2026-10-05)
# empty_response fix (2026-10-06, docs/empty-response.md): 13 of 119 runs ended with one turn that used the whole output limit on
# reasoning. Cap per turn 24000 (only 3 of 1997 turns were legitimately longer), one retry at another temperature (the same
# sample ran away again in 7 of 13 runs), and the stream guard (a streamed turn with only reasoning is cut after that many reasoning
# tokens) which stays off until it is verified on the cloud model (env STREAM_GUARD, for example 9000).
TEACHER_MAX_TOKENS = 24000
EMPTY_RETRIES = 1
RETRY_TEMPERATURE = 0.8
STREAM_GUARD = int(os.environ["STREAM_GUARD"]) if os.environ.get("STREAM_GUARD") else None
CDS_CALLS = 100
def new_agent():
return LlmAgent(MODEL, loop_guard=3, max_tokens=TEACHER_MAX_TOKENS, empty_retries=EMPTY_RETRIES,
retry_temperature=RETRY_TEMPERATURE, stream_guard=STREAM_GUARD) # tool-call budget for tasks with a CDS contract object (eval keeps 60; Kral 2026-10-05)
LOCK = threading.Lock()
@@ -74,7 +87,7 @@ def one(task_id, attempt, stop):
print("BUDGET", e, flush=True)
return
run_no = RUN_BASE + (int("".join(c for c in task_id if c.isdigit())) - 1000) * 3 + attempt
agent = LlmAgent(MODEL, loop_guard=3)
agent = new_agent()
runner = Runner(POOL, OUT)
t0 = time.time()
row = {"task": task_id, "attempt": attempt, "run": run_no, "model": MODEL}
@@ -97,7 +110,7 @@ def one(task_id, attempt, stop):
for d in glob.glob(os.path.join(OUT, f"{run_no}_{task_id}_*")):
os.makedirs(os.path.join(OUT, "_aborted"), exist_ok=True)
os.rename(d, os.path.join(OUT, "_aborted", os.path.basename(d) + "_" + str(int(time.time()))))
agent = LlmAgent(MODEL, loop_guard=3)
agent = new_agent()
score = (rep.get("score") or {}).get("total")
rec = os.path.exists(os.path.join(run_dir, "record.json"))
row.update(score=score, setup_failed=bool(rep.get("setup_failed")), end_reason=rep.get("end_reason"),

22
scripts_probe/after_owntests.sh Executable file
View File

@@ -0,0 +1,22 @@
#!/bin/sh
# Waits for harness.owntests to end, then: report, rebuild the stage 2 set with the own-test score as metadata, update the HF dataset, commit.
cd "$HOME/projects/abap-llm/harness" || exit 1
PID="$1"
while kill -0 "$PID" 2>/dev/null; do sleep 60; done
python3 train/own_test_report.py > runs/own_test_report.log 2>&1
train/.venv/bin/python train/build_stage2.py --out runs/stage2_data --hook hooks_example:own_test_weight > runs/build_after_owntests.log 2>&1
python3 train/build_doc.py >> runs/build_after_owntests.log 2>&1
set -a; . ./.env; set +a
train/.venv/bin/python - >> runs/build_after_owntests.log 2>&1 <<'PY'
import os
from huggingface_hub import HfApi
api = HfApi(token=os.environ["HF_TOKEN"])
for f in ("stage2_train.jsonl", "stage2_valid.jsonl", "stage2_reserve.jsonl", "build_report.json", "README.md"):
api.upload_file(path_or_fileobj="runs/stage2_data/" + f, path_in_repo=f, repo_id="erhankeseli/abap-stage2-data", repo_type="dataset")
print("uploaded")
PY
export GIT_AUTHOR_NAME=Kral GIT_AUTHOR_EMAIL=kral@local GIT_COMMITTER_NAME=Kral GIT_COMMITTER_EMAIL=kral@local
git add docs train && git commit -q -m "D: own-test mutation scores of the accepted trajectories (metadata only), stage 2 set rebuilt with the score
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>"
echo "after_owntests done $(date)" >> runs/own_test_report.log

View File

@@ -21,7 +21,7 @@ CLASS zprobe0eq_locks IMPLEMENTATION.
TABLES enq = lt_enq
EXCEPTIONS communication_failure = 1 system_failure = 2 OTHERS = 3.
DATA(lv_n) = 0.
LOOP AT lt_enq ASSIGNING FIELD-SYMBOL(<l>) WHERE garg CS 'ZPROBE0' OR garg CS 'Z4AJ' OR garg CS 'Z4AE'.
LOOP AT lt_enq ASSIGNING FIELD-SYMBOL(<l>).
lv_n = lv_n + 1.
DATA(lv_line) = ||.
DO.
@@ -35,15 +35,16 @@ CLASS zprobe0eq_locks IMPLEMENTATION.
ENDDO.
out->write( lv_line ).
ENDLOOP.
out->write( |enqueue entries matching the probe prefixes: { lv_n } of { lines( lt_enq ) } (subrc { lv_subrc })| ).
out->write( |enqueue entries listed: { lv_n } of { lines( lt_enq ) } (subrc { lv_subrc })| ).
ENDMETHOD.
ENDCLASS.
"""
def ensure_reader(m):
def ensure_reader(m, force=False):
e, t = m.call("sap_search_object", {"query": READER})
if READER not in t:
m.call("sap_create_object", {"objectType": "CLAS", "objectName": READER, "packageName": "$TMP", "description": "enqueue reader probe"})
if READER not in t or force:
if READER not in t:
m.call("sap_create_object", {"objectType": "CLAS", "objectName": READER, "packageName": "$TMP", "description": "enqueue reader probe"})
e, t = m.call("sap_push_source", {"objectType": "CLAS", "objectName": READER, "source": READER_SRC})
if '"success":true' not in t.replace(" ", ""):
print("reader activation:", t[:600])
@@ -54,5 +55,5 @@ def read_locks(m):
if __name__ == "__main__":
with McpClient() as m:
ensure_reader(m)
print(read_locks(m)[:2500])
ensure_reader(m, force=True)
print(read_locks(m)[:3500])

View File

@@ -180,3 +180,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- 2026-10-05 22:35 **pipeline stopped at 21:43 by an infrastructure outage, found 22:26 (my miss).** The MCP server (127.0.0.1:3000) refused connections for a short time; three runs got `URLError(ConnectionRefusedError)` and the pipeline counted three equal harness errors and stopped (rule). Not a model or harness bug: A4H was up (11 h), the MCP server answers again. New `harness/infra.py`: an outage of MCP or A4H is waited out (check every 30 s, two answers in a row), the run's objects are deleted, the same run starts again; it is not counted as an event or a failure (`runs/traj/outages.jsonl`). Same for series A (applies after a restart of `harness.localqwen`; the running series keeps the old code and would skip the task). The 3 interrupted runs were cleaned (11 objects) and get their second attempt. Budget: panel 50.84 at ledger 113.36 (2.54 ledger per usage in the last stretch); guard recomputed: panel left 9.16 - 3.0 reserve = 6.2 usage x 1.2 = 7.4 ledger, `BUDGET_LIMIT_USD` 129 (guard 121). Dashboard: a stopped pipeline shows no burn rate or ETA.
- 2026-10-05 22:40 cause of the 21:43 outage confirmed by Kral: he restarted Eclipse (the EPOD MCP server runs inside it). Three runs were writing at that moment; their objects could be deleted afterwards (no stale lock), so a restart of the MCP server in the middle of a write did not leak a lock in this case. Rule: restarting Eclipse is fine now (the pipeline waits and reruns), but check `python3 scripts_probe/lockprobe.py` for stale locks afterwards.
- 2026-10-06 06:40 **correction of the 22:40 entry and a new finding.** (1) An Eclipse restart in the middle of a write does leak a lock: `ZCL_Z4CGT1HJ_JOB_COST_TEST` (enqueue SEOCLSENQ, mode X, created 21:46, three minutes after the 21:43 restart) stayed locked until the next morning; details in `docs/epod-lock-leak.md` (case 3). I had missed it because the teardown only cleaned names that start with the run prefix. SM12 is needed (Kral). (2) **Leftover objects with the prefix inside the name** (ZCL_<prefix>_..., the model's own test classes): 27 stayed in A4H, and **57 of 90 accepted trajectories contain other runs' object names in tool results** (mostly `sap_inactive_objects`, in 30 cases the model pulled another run's class). Fixed: `harness/proxy.py` hides them (search, usage, inactive list) and answers "not available" for reads and writes of another run's object; `harness/runner.py` finds mid-name objects for teardown; `harness/sweep.py` deleted the leftovers (28 objects; the locked class remains). The 90 trajectories are not changed yet: see the data scrub option of `train/build_stage2.py`.

236
train/analyze.py Normal file
View File

@@ -0,0 +1,236 @@
"""Analysis of the accepted DeepSeek trajectories (Opus item B, 2026-10-06): repair taxonomy, near duplicates,
empty_response, and the behaviors the teacher shows and Qwen lacks.
python3 train/analyze.py writes runs/analysis/analysis.json and prints the numbers
"""
import glob
import json
import os
import re
import statistics
import sys
from collections import Counter
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
from harness import mix # noqa: E402
from harness.generator import RESERVED # noqa: E402
WRITE = ("sap_push_source", "sap_push_element", "sap_push_message", "sap_activate", "sap_create_object")
SEARCH_LIKE = ("sap_search_object", "sap_pull_source", "sap_object_members", "sap_object_structure", "sap_element_info",
"sap_usage_references", "sap_sql_query", "sap_inactive_objects")
SIG = re.compile(r"(?:IMPORTING|EXPORTING|RETURNING|CHANGING)[^.]*?\bTYPE\s+(?:c|n|p|x)\s+LENGTH\b|\bVALUE\([^)]*\)\s+TYPE\s+(?:c|n|p|x)\s+LENGTH\b", re.I | re.S)
def pct(v, p):
v = sorted(v)
return v[min(len(v) - 1, int(p * len(v)))] if v else None
def calls_of(messages):
"""[(index, tool name, args dict, result text)] in order."""
res = {m.get("tool_call_id"): m["content"] for m in messages if m["role"] == "tool"}
out = []
for i, m in enumerate(messages):
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
out.append((i, c["function"]["name"], a, res.get(c.get("id"), "")))
return out
def failed(tool, text):
t = (text or "").replace(" ", "")
return tool in WRITE and (t.startswith("ERROR:") or '"success":false' in t)
def error_text(text):
try:
o = json.loads(text.replace("ERROR: ", "", 1) if text.startswith("ERROR:") else text)
except ValueError:
return text[:200]
if not isinstance(o, dict):
return text[:200]
msgs = [m.get("message", "") for m in (o.get("activation") or {}).get("messages", []) if m.get("severity") == "E"]
out = " | ".join(msgs) or str(o.get("error") or o.get("message") or "")
sc = o.get("syntaxCheck")
if sc:
out += " | syntaxCheck: " + "; ".join(m.get("message", "") for m in sc.get("messages", []))
return out[:400]
def classify(src, err):
"""Labels of a failed write (a write can match more than one)."""
labels = []
if src and SIG.search(src):
labels.append("type_c_length_in_signature")
low = (err or "").lower()
if "longer than the allowed 30 characters" in low or "30 characters" in low:
labels.append("name_over_30")
words = {w for w in re.findall(r"[A-Za-z_]+", err or "")}
if "is not valid" in low or "reserved" in low or any(w.upper() in RESERVED and w.upper() in (err or "").upper() for w in words if len(w) > 3 and w.isupper()):
if re.search(r"\b(?:" + "|".join(sorted(RESERVED, key=len, reverse=True)) + r")\b", err or "", re.I) or "reserved" in low:
labels.append("reserved_word")
return labels
def norm(err):
e = re.sub(r"\bZ[0-9A-Z]{7}_\w+", "<OBJ>", err or "")
e = re.sub(r"\b[A-Z][A-Z0-9_]{5,}\b", "<NAME>", e)
e = re.sub(r"\d+", "N", e)
return e[:110]
def behaviors(messages):
cs = calls_of(messages)
out = {"first_write": None, "searches_before_first_write": 0, "max_consecutive_search_object": 0, "calls": len(cs),
"failed_writes": 0, "repaired": False, "calls_to_repair": None, "same_push_repeats": 0}
run_so = 0
first_fail = None
last_src = {}
for n, (i, tool, a, res) in enumerate(cs):
if tool == "sap_search_object":
run_so += 1
out["max_consecutive_search_object"] = max(out["max_consecutive_search_object"], run_so)
else:
run_so = 0
if tool in ("sap_push_source", "sap_push_element", "sap_push_message") and out["first_write"] is None:
out["first_write"] = n
if out["first_write"] is None and tool in SEARCH_LIKE:
out["searches_before_first_write"] += 1
if tool == "sap_push_source":
key = (a.get("objectName"), a.get("includeType"))
if last_src.get(key) == a.get("source") and a.get("source"):
out["same_push_repeats"] += 1
last_src[key] = a.get("source")
if failed(tool, res):
out["failed_writes"] += 1
if first_fail is None:
first_fail = (n, a.get("objectName"), a.get("source"))
elif first_fail and tool in ("sap_push_source", "sap_push_element") and '"success":true' in (res or "").replace(" ", "") \
and a.get("objectName") == first_fail[1] and a.get("source") != first_fail[2] and not out["repaired"]:
out["repaired"] = True
out["calls_to_repair"] = n - first_fail[0]
return out
def shingles(text, k=5):
t = re.sub(r"\bz[0-9a-z]{7}_", "", text.lower())
w = re.findall(r"\w+", t)
return {" ".join(w[i:i + k]) for i in range(max(len(w) - k + 1, 0))}
def final_sources(messages):
"""Last pushed source per (object, include): the trajectory's own code."""
last = {}
for i, tool, a, res in calls_of(messages):
if tool == "sap_push_source" and a.get("source") and '"success":true' in (res or "").replace(" ", ""):
last[(a.get("objectName"), a.get("includeType"))] = a["source"]
return "\n".join(last.values())
def main():
rows = [json.loads(l) for l in open(os.path.join(ROOT, "runs", "traj", "summary.jsonl"))]
acc_recs, empty_runs, all_turn_out = [], [], []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
if not os.path.exists(p):
continue
rec = json.load(open(p))
ok, why = acc.judge(rec, r, 80)
md = rec["metadata"]
if ok:
acc_recs.append((r, rec))
all_turn_out += [u.get("completion_tokens", 0) for u in md.get("turn_usage", [])]
if md.get("end_reason") == "empty_response":
empty_runs.append((r, rec))
out = {"accepted": len(acc_recs), "runs": len(rows)}
# ---- repair taxonomy
fail_ct, labels_ct, traj_with, repaired_with, top = 0, Counter(), Counter(), Counter(), Counter()
per_label_examples = {}
kinds = Counter()
for r, rec in acc_recs:
seen = {}
for i, tool, a, res in calls_of(rec["messages"]):
if failed(tool, res):
fail_ct += 1
err = error_text(res)
top[norm(err)] += 1
for l in classify(a.get("source"), err):
labels_ct[l] += 1
seen[l] = True
per_label_examples.setdefault(l, (r["task"], err[:160]))
b = behaviors(rec["messages"])
for l in seen:
traj_with[l] += 1
if b["repaired"]:
repaired_with[l] += 1
out["repair_taxonomy"] = {"failed_writes_in_accepted": fail_ct, "step3_errors": {
l: {"failed_writes": labels_ct[l], "trajectories": traj_with[l], "trajectories_with_repair": repaired_with[l],
"example": per_label_examples.get(l)} for l in ("type_c_length_in_signature", "reserved_word", "name_over_30")},
"top_error_messages": top.most_common(15)}
# ---- teacher behaviors vs Qwen
def stats(recs):
b = [behaviors(x["messages"]) for x in recs]
had_fail = [x for x in b if x["failed_writes"]]
sb = [x["searches_before_first_write"] for x in b if x["first_write"] is not None]
return {"n": len(b), "with_a_failed_write": len(had_fail),
"repair_after_first_error": sum(x["repaired"] for x in had_fail), "median_calls_to_repair":
statistics.median([x["calls_to_repair"] for x in had_fail if x["calls_to_repair"] is not None] or [None]),
"median_searches_before_first_write": statistics.median(sb) if sb else None, "p90_searches_before_first_write": pct(sb, .9),
"max_consecutive_search_object": max([x["max_consecutive_search_object"] for x in b] or [0]),
"runs_with_no_write": sum(1 for x in b if x["first_write"] is None),
"runs_with_same_push_repeat": sum(1 for x in b if x["same_push_repeats"] > 0)}
out["teacher"] = stats([rec for _, rec in acc_recs])
qrecs = []
for p in glob.glob(os.path.join(ROOT, "runs", "local_qwen", "runs", "*", "record.json")):
qrecs.append(json.load(open(p)))
out["qwen_series_A"] = stats(qrecs)
out["qwen_series_A"]["note"] = "all 20 runs of series A (accepted and failed)"
# ---- near duplicates
sh = [(r["task"], r["run"], shingles(final_sources(rec["messages"])), shingles(rec["messages"][-1].get("content") or "")) for r, rec in acc_recs]
dup = []
for i in range(len(sh)):
for j in range(i + 1, len(sh)):
a, b = sh[i], sh[j]
for k, nm in ((2, "code"), (3, "report")):
if a[k] and b[k]:
jac = len(a[k] & b[k]) / len(a[k] | b[k])
if jac >= (0.8 if nm == "code" else 0.85):
dup.append({"a": f"{a[0]}_r{a[1]}", "b": f"{b[0]}_r{b[1]}", "what": nm, "jaccard": round(jac, 2), "same_task": a[0] == b[0]})
out["near_duplicates"] = {"pairs_flagged": len(dup), "same_task_pairs": sum(1 for d in dup if d["same_task"]), "examples": dup[:12]}
# ---- empty_response
er = []
for r, rec in empty_runs:
md = rec["metadata"]
tu = md.get("turn_usage", [])
last = [u.get("completion_tokens", 0) for u in tu[-3:]]
cs = calls_of(rec["messages"])
lastcall = cs[-1] if cs else None
er.append({"task": r["task"], "kind": mix.kind_of_task_dir(r["task"]), "category": rec["task"]["category"], "calls": md.get("tool_calls"),
"turns": len(tu), "last3_output_tokens": last, "last_tool": lastcall[1] if lastcall else None,
"last_tool_result_chars": len(lastcall[3]) if lastcall else None,
"context_tokens_last": (tu[-1].get("prompt_tokens") if tu else None)})
out["empty_response"] = {"runs": len(er), "of": len(rows), "by_kind": Counter(e["kind"] for e in er),
"by_last_tool": Counter(e["last_tool"] for e in er), "detail": er,
"output_tokens_of_normal_turns": {"p50": pct(all_turn_out, .5), "p90": pct(all_turn_out, .9),
"p99": pct(all_turn_out, .99), "p999": pct(all_turn_out, .999), "max": max(all_turn_out or [0]),
"turns": len(all_turn_out), "turns_over_16k": sum(1 for x in all_turn_out if x > 16000),
"turns_over_20k": sum(1 for x in all_turn_out if x > 20000)}}
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "analysis.json"), "w"), indent=1, default=dict)
print(json.dumps(out, indent=1, default=dict)[:6500])
if __name__ == "__main__":
main()

70
train/build_doc.py Normal file
View File

@@ -0,0 +1,70 @@
"""Writes docs/stage2-build.md from runs/stage2_data/build_report.json (run after train/build_stage2.py)."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
rep = json.load(open(os.path.join(ROOT, "runs", "stage2_data", "build_report.json")))
tr, va, rs = rep["train"], rep["valid"], rep["reserve"]
s1 = rep["stage1_stats"]["train"]["tokens"]
s1d = rep["stage1_stats"]["train"]["docs"]
L2 = tr["loss_tokens"]
def row(e1, e2, L=L2):
a, b = s1 * e1, L * e2
return f"| {e1} | {e2} | {a / 1e6:.2f} M | {b / 1e6:.2f} M | {100 * b / (a + b):.0f} % | {(a + b) / 1e6:.1f} M |"
drop = rep["dropped"]
dropped_lines = "\n".join(f"- {k}: {v}" for k, v in drop.items())
kinds = "\n".join(f"| {k} | {v['samples']} | {v['families']} | {v['short']} |" for k, v in rep["kinds"].items())
doc = f"""# Stage 2 training set, build of {rep['built']}
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
## Result
{rep['steps']['accepted_trajectories']} accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> {rep['steps']['after_eval_overlap']} after the eval overlap check
-> {rep['steps']['after_token_limit']} after the 48k limit (none cut) -> CLAS cap 35 %: {rep['steps']['clas_cap']['clas_kept']} CLAS kept, **{rep['steps']['clas_cap']['clas_reserve']} CLAS in reserve** (`stage2_reserve.jsonl`)
-> **{tr['samples']} train + {va['samples']} valid samples** ({tr['tokens'] / 1e6:.2f} M + {va['tokens'] / 1e6:.2f} M tokens, {tr['loss_tokens'] / 1e6:.2f} M loss tokens in train, p50 {tr['p50']}, p95 {tr['p95']}, max {tr['max']}, repair share {tr['repair_share']}).
Dropped:
{dropped_lines}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
## Not good enough yet (kinds under the minimum of {rep['settings']['min_per_kind']})
| kind | samples | families | short |
|---|---|---|---|
{kinds}
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides)
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
{row(2, 3)}
{row(1, 3)}
{row(0.5, 3)}
{row(0.33, 3)}
{row(1, 3, 3 * L2)}
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
"""
open(os.path.join(ROOT, "docs", "stage2-build.md"), "w").write(doc)
print("written", len(doc))

320
train/build_stage2.py Normal file
View File

@@ -0,0 +1,320 @@
"""Stage 2 training set builder (Opus item C, 2026-10-06).
train/.venv/bin/python train/build_stage2.py [--out runs/stage2_data] [--max-tokens 48000] [--clas-cap 0.35]
[--min-per-kind 25] [--valid-frac 0.10] [--hook module:function]
Input: runs/traj (DeepSeek trajectories; the local Qwen series is never read). Steps: acceptance filter (score, end reason, no
harness text, loop repeats trimmed) -> eval overlap check -> Qwen 3.8 chat template (thinking off) with the tokenizer round trip
check -> samples over --max-tokens are dropped (never cut) -> CLAS capped at --clas-cap of the stage 2 samples (the rest goes to
`stage2_reserve.jsonl`) -> shortage report for kinds under --min-per-kind -> validation split by task family (K variants and all
trajectories of a task on the same side) -> optional hook (own-test score, weights) -> files + data card.
Loss mask: `assistant_spans` are character spans of the assistant turns (tool calls and the final report, with <|im_end|>);
nothing of the system turn (tool schemas), the user turn or the tool results is in a span. `assistant_tokens` counts the loss tokens.
"""
import argparse
import importlib
import re
import json
import os
import random
import statistics
import sys
import time
from collections import Counter, defaultdict
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
import to_qwen as tq # noqa: E402
from harness import mix, overlap # noqa: E402
from transformers import AutoTokenizer # noqa: E402
POOL = os.path.join(ROOT, "tasks_gen", "train")
START_PREFIX = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
LIST_TOOLS = ("sap_inactive_objects", "sap_search_object", "sap_usage_references")
READ_TOOLS = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
def foreign_name(name, own):
n = (name or "").upper()
m = START_PREFIX.match(n) or MID_PREFIX.match(n)
return bool(m) and m.group(1).upper() != own.upper()
def scrub_foreign(msgs, own_prefix):
"""Remove other runs' leftover objects from list results (the proxy hides them since 2026-10-06; the older records still have them).
Returns (messages, removed list entries, number of reads of a foreign object)."""
tool_of = {}
for m in msgs:
if m["role"] == "assistant":
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
tool_of[c.get("id")] = (c["function"]["name"], a)
removed = reads = 0
out = []
for m in msgs:
if m["role"] == "tool":
name, args = tool_of.get(m.get("tool_call_id"), ("", {}))
if name in READ_TOOLS and foreign_name(args.get("objectName"), own_prefix):
reads += 1
if name in LIST_TOOLS:
body = m["content"]
prefix = "ERROR: " if body.startswith("ERROR: ") else ""
try:
data = json.loads(body[len(prefix):])
except ValueError:
data = None
if isinstance(data, list):
keep = [d for d in data if not (isinstance(d, dict) and foreign_name(d.get("name"), own_prefix))]
if len(keep) != len(data):
removed += len(data) - len(keep)
m = dict(m, content=prefix + json.dumps(keep))
out.append(m)
return out, removed, reads
def pct(v, p):
v = sorted(v)
return v[min(len(v) - 1, int(p * len(v)))] if v else None
def family_of(task_id):
try:
t = json.load(open(os.path.join(POOL, task_id, "task.json")))
except OSError:
return task_id
return t.get("base_task") or task_id
def loss_tokens(tok, text, spans):
enc = tok(text, add_special_tokens=False, return_offsets_mapping=True)
n = 0
for (a, b) in enc["offset_mapping"]:
if any(a >= s and b <= e for s, e in spans):
n += 1
return len(enc["input_ids"]), n
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--traj", default=os.path.join(ROOT, "runs", "traj"))
ap.add_argument("--out", default=os.path.join(ROOT, "runs", "stage2_data"))
ap.add_argument("--max-tokens", type=int, default=48000)
ap.add_argument("--clas-cap", type=float, default=0.35)
ap.add_argument("--min-per-kind", type=int, default=25)
ap.add_argument("--min-score", type=float, default=80)
ap.add_argument("--valid-frac", type=float, default=0.10)
ap.add_argument("--seed", type=int, default=20261006)
ap.add_argument("--no-scrub", action="store_true", help="keep other runs' leftover objects in list results (default: removed)")
ap.add_argument("--keep-foreign-reads", action="store_true", help="keep trajectories in which the model read another run's object (default: dropped)")
ap.add_argument("--hook", help="module:function; function(row) -> None to drop, or a float weight (repeat factor), or a dict "
"{'keep': bool, 'weight': float, 'extra': {...}} (for example the own-test mutation score)")
a = ap.parse_args()
os.makedirs(a.out, exist_ok=True)
hook = None
if a.hook:
mod, fn = a.hook.split(":")
sys.path.insert(0, os.path.join(ROOT, "train"))
hook = getattr(importlib.import_module(mod), fn)
tok = AutoTokenizer.from_pretrained(os.path.expanduser("~/models/Qwen3.8-27B-4bit"))
evals = overlap.load_pool("eval")
report = {"built": time.strftime("%F %T"), "settings": vars(a), "dropped": defaultdict(Counter), "steps": {}}
# 1 accepted trajectories
rows = [json.loads(l) for l in open(os.path.join(a.traj, "summary.jsonl"))]
cand = []
for r in rows:
p = os.path.join(a.traj, r.get("run_dir") or "-", "record.json")
if not os.path.exists(p):
continue
rec = json.load(open(p))
ok, why = acc.judge(rec, r, a.min_score)
if not ok:
continue
msgs, removed = acc.trim_loops(rec["messages"])
scrubbed = freads = 0
if not a.no_scrub:
msgs, scrubbed, freads = scrub_foreign(msgs, rec["prefix"])
if freads and not a.keep_foreign_reads:
kind0 = mix.kind_of_task_dir(r["task"])
report["dropped"][kind0]["read_of_another_runs_object"] += 1
continue
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs), "scrubbed": scrubbed, "freads": freads})
report["steps"]["accepted_trajectories"] = len(cand)
# 2 eval overlap at build time (the task against every eval task)
kept = []
for c in cand:
tid = c["row"]["task"]
kind = mix.kind_of_task_dir(tid)
try:
hits = overlap.check(overlap.load_task(os.path.join(POOL, tid)), evals)
except (OSError, ValueError):
hits = []
if hits:
report["dropped"][kind]["eval_overlap"] += 1
continue
c["kind"] = kind
kept.append(c)
report["steps"]["after_eval_overlap"] = len(kept)
# 3 Qwen template + tokens + loss tokens; drop over the limit (never cut)
samples = []
for c in kept:
msgs = tq.to_template_messages(c["msgs"])
tools = c["rec"]["tools"]
text = tok.apply_chat_template(msgs, tools=tools, tokenize=False, enable_thinking=False)
prefix = tok.apply_chat_template(msgs[:2], tools=tools, tokenize=False, add_generation_prompt=True, enable_thinking=False)
if not text.startswith(prefix) or not tq.check_calls(msgs, text):
report["dropped"][c["kind"]]["template_roundtrip"] += 1
continue
spans = tq.spans(text)
n, nl = loss_tokens(tok, text, spans)
if n > a.max_tokens:
report["dropped"][c["kind"]]["over_%dk" % (a.max_tokens // 1000)] += 1
continue
t = c["rec"]["task"]
samples.append({"id": f"{c['row']['task']}_r{c['row']['run']}", "task": c["row"]["task"], "family": family_of(c["row"]["task"]),
"kind": c["kind"], "category": t.get("category"), "object_type": t.get("object_type"),
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"], "foreign_entries_removed": c["scrubbed"],
"score": c["rec"]["metadata"].get("score"), "teacher": c["rec"]["teacher"], "weight": 1.0,
"n_tokens": n, "assistant_tokens": nl, "text": text, "assistant_spans": spans})
report["steps"]["after_token_limit"] = len(samples)
# 4 CLAS cap: keep the CLAS samples with a repair and the highest score first; the rest is the reserve
by_kind = defaultdict(list)
for s in samples:
by_kind[s["kind"]].append(s)
non_clas = sum(len(v) for k, v in by_kind.items() if k != "CLAS")
clas = sorted(by_kind.get("CLAS", []), key=lambda s: (not s["repair"], -(s["score"] or 0), s["id"]))
limit = int(a.clas_cap / (1 - a.clas_cap) * non_clas) if non_clas else len(clas)
reserve = clas[limit:]
by_kind["CLAS"] = clas[:limit]
report["steps"]["clas_cap"] = {"clas_before": len(clas), "clas_kept": len(by_kind["CLAS"]), "clas_reserve": len(reserve), "limit": limit}
pool = [s for v in by_kind.values() for s in v]
# 5 shortage report (never filled by copies)
report["kinds"] = {}
for k in mix.TYPE_SHARE:
n = len(by_kind.get(k, []))
report["kinds"][k] = {"samples": n, "min_required": a.min_per_kind, "short": max(0, a.min_per_kind - n),
"families": len({s["family"] for s in by_kind.get(k, [])})}
# 6 hook (own-test score, weights)
if hook:
out = []
for s in pool:
r = hook(s)
if r is None or r is False:
report["dropped"][s["kind"]]["hook"] += 1
continue
if isinstance(r, dict):
if not r.get("keep", True):
report["dropped"][s["kind"]]["hook"] += 1
continue
s["weight"] = float(r.get("weight", 1.0))
s.update(r.get("extra") or {})
elif isinstance(r, (int, float)) and not isinstance(r, bool):
s["weight"] = float(r)
out.append(s)
pool = out
# 7 validation split by family, stratified by kind
rnd = random.Random(a.seed)
fam_kind = {}
for s in pool:
fam_kind.setdefault(s["family"], s["kind"])
valid_fam = set()
for k in mix.TYPE_SHARE:
fams = sorted(f for f, kk in fam_kind.items() if kk == k)
rnd.shuffle(fams)
nv = round(len(fams) * a.valid_frac)
if len(fams) >= 4:
nv = max(nv, 1)
valid_fam |= set(fams[:nv])
train = [s for s in pool if s["family"] not in valid_fam]
valid = [s for s in pool if s["family"] in valid_fam]
for name, data in (("stage2_train", train), ("stage2_valid", valid), ("stage2_reserve", reserve)):
with open(os.path.join(a.out, name + ".jsonl"), "w") as f:
for s in data:
f.write(json.dumps(s) + "\n")
# 8 report and data card
def summary(data):
toks = [s["n_tokens"] for s in data]
return {"samples": len(data), "tokens": sum(toks), "loss_tokens": sum(s["assistant_tokens"] for s in data),
"p50": pct(toks, .5), "p90": pct(toks, .9), "p95": pct(toks, .95), "max": max(toks or [0]),
"by_kind": dict(Counter(s["kind"] for s in data)), "by_category": dict(Counter(s["category"] for s in data)),
"repair_share": round(sum(s["repair"] for s in data) / max(len(data), 1), 2)}
report["train"], report["valid"], report["reserve"] = summary(train), summary(valid), summary(reserve)
report["dropped"] = {k: dict(v) for k, v in report["dropped"].items()}
s1 = json.load(open(os.path.join(ROOT, "train", "data", "stats.json"))) if os.path.exists(os.path.join(ROOT, "train", "data", "stats.json")) else {}
report["stage1_stats"] = s1
json.dump(report, open(os.path.join(a.out, "build_report.json"), "w"), indent=1, default=str)
open(os.path.join(a.out, "README.md"), "w").write(data_card(report))
print(json.dumps({k: report[k] for k in ("steps", "train", "valid", "reserve", "dropped")}, indent=1, default=str)[:3500])
print("short kinds:", {k: v["short"] for k, v in report["kinds"].items() if v["short"]})
def data_card(r):
t, v = r["train"], r["valid"]
kinds = "\n".join(f"| {k} | {x['samples']} | {x['families']} | {x['short'] or ''} |" for k, x in r["kinds"].items())
drop = "\n".join(f"- {k}: {d}" for k, d in r["dropped"].items()) or "- nothing dropped"
return f"""---
license: mit
task_categories: [text-generation]
tags: [abap, sap, agentic, tool-use, sft]
private: true
---
# ABAP stage 2 agent trajectories (built {r['built']})
Tool-using ABAP development trajectories for supervised fine-tuning of Qwen 3.8 27B. A teacher model solved generated ABAP tasks on a real
SAP ABAP Platform system (A4H, SAP_BASIS 816) through ADT tools; only runs that passed the harness gates and scored at least {r['settings']['min_score']:.0f} of 100 are kept.
## Sources and licenses
- Teacher: DeepSeek V4.1 Flash (MIT license), via Ollama cloud. No output of Claude or other restricted models.
- Tasks: generated by the same teacher (spec, seed objects, hidden ABAP Unit tests, reference), validated on A4H (oracle 100, null 0, mutation check). Eval tasks are not in this data (overlap check at build time).
- The local Qwen runs (series A) are not in this data.
## Format
One JSON per line: `text` (the whole conversation in the Qwen 3.8 chat template, thinking off, Qwen XML tool call format, the 20 tool schemas in the system turn),
`assistant_spans` (character spans that carry the loss: assistant turns with tool calls and the final report; none of the system turn, user turn or tool results),
`n_tokens`, `assistant_tokens`, `kind`, `category`, `object_type`, `family` (task family: a K variant has the family of its base task), `repair` (an error followed by a fix), `score`, `weight`.
## Size
| split | samples | tokens | loss tokens | p50 | p90 | p95 | max | repair share |
|---|---|---|---|---|---|---|---|---|
| train | {t['samples']} | {t['tokens']} | {t['loss_tokens']} | {t['p50']} | {t['p90']} | {t['p95']} | {t['max']} | {t['repair_share']} |
| valid | {v['samples']} | {v['tokens']} | {v['loss_tokens']} | {v['p50']} | {v['p90']} | {v['p95']} | {v['max']} | {v['repair_share']} |
Samples over {r['settings']['max_tokens']} tokens are dropped, never cut. CLAS is capped at {int(100 * r['settings']['clas_cap'])} % of the samples (the rest is in `stage2_reserve.jsonl`).
## Kinds (target share: CLAS 28, INTF 7, CDS 25, FUNC 15, PROG 10, TABL 8, STRU 2, MSAG 2.5, exception 2.5 percent)
| kind | samples | families | short of the minimum {r['settings']['min_per_kind']} |
|---|---|---|---|
{kinds}
## Dropped
{drop}
## Known limits
- Small data (see the table); kinds marked short have fewer than the minimum samples.
- Until 2026-10-06 the test system still held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes). Their names are removed from list results (`sap_inactive_objects`, searches) in these samples; trajectories in which the model read such an object are dropped.
- The proxy added local abaplint messages (`syntaxCheck`) to some failed writes; the EPOD server does not do this itself.
- The teacher sees only the tool results of its own runs; a run that repaired an error is kept with its error (60 to 65 % of the samples).
- Validation split is by task family; do not mix valid samples into the training set.
"""
if __name__ == "__main__":
main()

163
train/hf_train_bf16.py Normal file
View File

@@ -0,0 +1,163 @@
# /// script
# requires-python = ">=3.10"
# dependencies = ["unsloth", "datasets", "transformers", "huggingface_hub"]
# ///
"""Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go.
Memory test first (a few dollars, finds the largest sequence length that fits, no data needed):
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000
Real run (after the memory test and Kral's decision on the ratio):
hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s1-epochs 1 --s2-epochs 3 --out erhankeseli/abap-mixed-adapter
Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns;
none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut.
Base model in bf16 (no nf4): the adapter then fits the MLX 4-bit base better (stage 1 finding, train/STATE.md).
"""
import argparse
import json
import os
import random
import time
def tokenize_masked(tok, text, spans, max_len):
"""input_ids and labels; labels are -100 outside the spans (spans=None: loss on every token). None when longer than max_len."""
enc = tok(text, add_special_tokens=False, return_offsets_mapping=True)
ids = enc["input_ids"]
if len(ids) > max_len:
return None
if spans is None:
return ids, list(ids)
labels = []
for tid, (a, b) in zip(ids, enc["offset_mapping"]):
labels.append(tid if any(a >= s and b <= e for s, e in spans) else -100)
return ids, labels
def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
rnd = random.Random(seed)
rows = []
for ep in range(int(s1_epochs)):
rows += [("s1", r["text"], None, 1.0) for r in stage1]
frac = s1_epochs - int(s1_epochs)
if frac:
rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))]
for ep in range(int(s2_epochs)):
for r in stage2:
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * max(1, round(float(r.get("weight", 1.0))))
rnd.shuffle(rows)
out, skipped = [], {"s1": 0, "s2": 0}
for src, text, spans, _ in rows:
t = tokenize_masked(tok, text, spans, max_len)
if t is None:
skipped[src] += 1
continue
out.append({"input_ids": t[0], "labels": t[1], "src": src})
return out, skipped
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--model", default="Qwen/Qwen3.8-27B")
ap.add_argument("--stage1", default="erhankeseli/abap-stage1-data")
ap.add_argument("--stage2", default="erhankeseli/abap-stage2-data")
ap.add_argument("--s1-file", default="train.jsonl")
ap.add_argument("--s2-file", default="stage2_train.jsonl")
ap.add_argument("--s2-valid", default="stage2_valid.jsonl")
ap.add_argument("--s1-epochs", type=float, default=1.0)
ap.add_argument("--s2-epochs", type=float, default=3.0)
ap.add_argument("--max-seq", type=int, default=48000)
ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--alpha", type=int, default=32)
ap.add_argument("--lr", type=float, default=5e-5)
ap.add_argument("--grad-accum", type=int, default=4)
ap.add_argument("--warmup-steps", type=int, default=20)
ap.add_argument("--no-offload", action="store_true", help="gradient checkpointing on the GPU instead of the Unsloth offload")
ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter")
ap.add_argument("--save-every", type=int, default=0)
ap.add_argument("--memory-test", action="store_true")
ap.add_argument("--sweep", default="16000,32000,48000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length")
a = ap.parse_args()
import torch
from unsloth import FastLanguageModel
from huggingface_hub import HfApi, hf_hub_download
token = os.environ["HF_TOKEN"]
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=a.max_seq, load_in_4bit=False, dtype=torch.bfloat16, token=token)
model = FastLanguageModel.get_peft_model(
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"],
use_gradient_checkpointing=True if a.no_offload else "unsloth", random_state=20261006)
print("GPU", torch.cuda.get_device_name(0), "x", torch.cuda.device_count(), "| weights allocated GB", round(torch.cuda.memory_allocated() / 2**30, 1), flush=True)
if a.memory_test:
text = "CLASS zcl_demo DEFINITION PUBLIC. METHODS run IMPORTING iv TYPE i. ENDCLASS.\n" * 4000
ids_all = tok(text, add_special_tokens=False)["input_ids"]
opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=1e-5)
model.train()
result = {"gpu": torch.cuda.get_device_name(0), "gpus": torch.cuda.device_count(), "offload": not a.no_offload, "sweep": []}
for n in [int(x) for x in a.sweep.split(",")]:
ids = torch.tensor([(ids_all * (n // len(ids_all) + 1))[:n]], device="cuda")
labels = ids.clone()
labels[:, : int(n * 0.68)] = -100 # about 32 % loss tokens, as in the stage 2 data
torch.cuda.reset_peak_memory_stats()
try:
times = []
for _ in range(a.steps):
t0 = time.time()
loss = model(input_ids=ids, labels=labels).loss
loss.backward()
opt.step()
opt.zero_grad(set_to_none=True)
torch.cuda.synchronize()
times.append(round(time.time() - t0, 1))
row = {"seq_len": n, "ok": True, "peak_gb": round(torch.cuda.max_memory_allocated() / 2**30, 1),
"reserved_gb": round(torch.cuda.max_memory_reserved() / 2**30, 1), "step_seconds": times, "loss": float(loss)}
except torch.cuda.OutOfMemoryError as e:
row = {"seq_len": n, "ok": False, "error": "OOM", "peak_gb": round(torch.cuda.max_memory_allocated() / 2**30, 1)}
result["sweep"].append(row)
print("MEMTEST", json.dumps(row), flush=True)
break
result["sweep"].append(row)
print("MEMTEST", json.dumps(row), flush=True)
torch.cuda.empty_cache()
print("RESULT", json.dumps(result), flush=True)
HfApi(token=token).upload_file(path_or_fileobj=json.dumps(result, indent=1).encode(), path_in_repo="memtest_%s_%dx.json" % (
result["gpu"].replace(" ", "_"), result["gpus"]), repo_id=a.out, repo_type="model", create_pr=False) if a.out else None
return
def rows(repo, fname):
p = hf_hub_download(repo, fname, repo_type="dataset", token=token)
return [json.loads(l) for l in open(p)]
s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file)
ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006)
loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex)
print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens",
round(sum(sum(1 for x in e["labels"] if x != -100) for e in ex if e["src"] == "s2") / max(loss_tok, 1), 2), flush=True)
from datasets import Dataset
from transformers import Trainer, TrainingArguments
class Collate:
def __call__(self, batch): # batch size 1; no padding needed
b = batch[0]
return {"input_ids": torch.tensor([b["input_ids"]]), "labels": torch.tensor([b["labels"]]),
"attention_mask": torch.ones(1, len(b["input_ids"]), dtype=torch.long)}
args = TrainingArguments(output_dir="out", per_device_train_batch_size=1, gradient_accumulation_steps=a.grad_accum, num_train_epochs=1,
learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=a.warmup_steps, optim="adamw_8bit", bf16=True,
logging_steps=1, save_strategy="no", report_to="none", seed=20261006, remove_unused_columns=False)
tr = Trainer(model=model, args=args, train_dataset=Dataset.from_list(ex), data_collator=Collate())
t0 = time.time()
tr.train()
info = {"examples": len(ex), "skipped": skipped, "train_seconds": round(time.time() - t0, 1), "rank": a.rank, "alpha": a.alpha, "lr": a.lr,
"peak_gpu_gb": round(torch.cuda.max_memory_allocated() / 2**30, 1), "s1_epochs": a.s1_epochs, "s2_epochs": a.s2_epochs}
print("RESULT", json.dumps(info), flush=True)
model.save_pretrained("adapter")
json.dump(info, open("adapter/job_result.json", "w"))
HfApi(token=token).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model")
print("PUSHED", a.out)
if __name__ == "__main__":
main()

28
train/hooks_example.py Normal file
View File

@@ -0,0 +1,28 @@
"""Hooks for train/build_stage2.py (--hook hooks_example:own_test_weight). A hook gets one sample (a dict) and returns
None / False (drop), a number (repeat weight), or {"keep": bool, "weight": float, "extra": {...}}."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
def identity(row):
return 1.0
def own_test_weight(row):
"""Item D (own-test mutation score, metadata only for now): reads runs/traj/<run>/own_test_mutation.json when it exists,
stores it as extra data and does NOT drop or reweight (Kral + Opus 2026-10-06: do not change the acceptance yet)."""
run = row["id"].split("_r")[-1]
for d in os.listdir(os.path.join(ROOT, "runs", "traj")):
if d.startswith(run + "_"):
p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json")
if os.path.exists(p):
m = json.load(open(p))
return {"keep": True, "weight": 1.0, "extra": {"own_test_mutation": m.get("score"), "own_test_mutants": m.get("mutants")}}
return {"keep": True, "weight": 1.0}
def repair_up(row):
"""Example of a weight: a trajectory with a repair counts twice."""
return 2.0 if row.get("repair") else 1.0

42
train/own_test_report.py Normal file
View File

@@ -0,0 +1,42 @@
"""Summary of the own-test mutation scores (item D): docs/own-test-mutation.md. Run after `python3 -m harness.owntests`."""
import glob
import json
import os
import statistics
from collections import Counter, defaultdict
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
res = [json.load(open(p)) for p in glob.glob(os.path.join(ROOT, "runs", "traj", "*", "own_test_mutation.json"))]
by = Counter(r.get("status") for r in res)
scored = [r for r in res if r.get("status") == "scored"]
ref_ok = [r for r in scored if r.get("tests_pass_on_reference")]
per_kind = defaultdict(list)
for r in ref_ok:
if r.get("score") is not None:
per_kind[r["kind"]].append(r["score"])
sc = [r["score"] for r in ref_ok if r.get("score") is not None]
bins = Counter("1.0" if s == 1 else ">=0.75" if s >= 0.75 else ">=0.5" if s >= 0.5 else "<0.5" for s in sc)
low = sorted([r for r in ref_ok if r.get("score") is not None and r["score"] < 0.75], key=lambda r: r["score"])
doc = f"""# Own-test mutation score of the accepted trajectories (item D, metadata only)
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
## Result
{len(res)} accepted trajectories looked at: {dict(by)}.
- Scored: {len(scored)}; the model's tests also passed on the correct reference: {len(ref_ok)} (the others are not reliable: a test that fails on the correct solution kills every mutant).
- Mean score over the reliable ones: {statistics.mean(sc):.2f}; median {statistics.median(sc):.2f}; distribution {dict(bins)} (n = {len(sc)}).
- By kind (mean, n): {{{", ".join(f"{k}: {statistics.mean(v):.2f} ({len(v)})" for k, v in sorted(per_kind.items()))}}}
- `no_own_tests`: {by.get('no_own_tests', 0)} trajectories were accepted without any own test (at most 85 points); `no_mutants`: {by.get('no_mutants', 0)}; `not_supported` (PROG: tests are inside the program): {by.get('not_supported', 0)}.
## Weakest (score under 0.75, tests pass on the reference)
""" + "\n".join(f"- {r['task']} run {r['run']} {r['kind']}: score {r['score']} ({r['killed']} of {r['valid']} valid mutants killed)" for r in low[:15]) + """
## Use
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).
"""
open(os.path.join(ROOT, "docs", "own-test-mutation.md"), "w").write(doc)
print(doc[:1500])