Mutation in generator; category H (stop) scoring and judge; category K variants and tool schema variant; evalset plan

- mutation.py: negation mutants, skip WHILE/DO blocks (endless loop blocked RFC ~8 min)
- generator: accept only after mutation check; survivors go back as repair feedback
- runner/judge.py: stop tasks scored 100/30/0; gap by keywords, else judge model
- proxy: tool schema variant generic_v0 (draft); generator make_k_variant
- evalset.py: 88 slots for step F (A-I, H, K)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-03 05:37:20 +02:00
parent 32f03fcad4
commit 0ba25faeea
20 changed files with 1303 additions and 22 deletions

View File

@@ -112,5 +112,7 @@ python3 -c "from harness.ledger import spent; print(spent())"
- When a write fails, run a syntax check and return its messages. Now the server says only "save failed"
(for example for `TYPE c LENGTH n` in a method signature). The model cannot see the cause and cannot repair;
this hurts eval and training. Workaround in the generator: local abaplint parser check.
- Unit test timeout: an endless loop in tested code blocks the one RFC connection for ~8 min
(mutation run 5201, G0002). Request: stop a unit test run after N seconds (DURATION SHORT = 60 s).
- Concurrency: queue calls per RFC connection, or use a connection pool.
- BDEF creation (needed for RAP tasks).