Mutation check done on all accepted tasks; mutant fixes; dump filter in proxy

- skip sy-subrc lines and CDS type lengths; CDS literal mutants; mutate helper classes
- proxy: sap_short_dumps shows only dumps of the current run
- 26/26 tasks pass; results in docs/faz1-tasarim.md 11g

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 06:02:59 +02:00
parent 0ba25faeea
commit f907351c0d
108 changed files with 7384 additions and 8 deletions

View File

@@ -96,12 +96,16 @@ python3 -c "from harness.ledger import spent; print(spent())"
task, first-attempt 1/5 (`docs/faz1-tasarim.md` 11f).
- Generator: static checks, local abaplint parser check, dependency order of seed/reference, budget floor,
max_tokens 80k. Runner: G2 for CDS with parameters, G6 ignores unknown standard superclasses.
- Step D: `harness/mutation.py` (deterministic mutants of the reference, hidden tests must fail).
- Step D done: `harness/mutation.py`, part of generator acceptance. All 26 tasks pass (`docs/faz1-tasarim.md` 11g).
Killed mutants are kept in `<task>/faulty/`.
- Step 5 code: H stop scoring (`judge.py`), K variants (`make_k_variant`), proxy tool schema `generic_v0` (draft).
Not yet run on SAP.
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
## 7. Next steps
1. Step D: run the mutation check on all accepted tasks; integrate it into the generator.
2. Use killed mutants as faulty references for own-test scoring (component (a)).
1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`).
2. Own-test scoring part (a): run the model's own tests against `faulty/` mutants.
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
checklist, Kral spot-checks ~10 flagged tasks.
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).