Hard batch G0022-G0026: 5/5 accepted; budget floor, CDS checks, mutation check
- generator: budget floor (2x oracle activations, 3x calls), static checks for CDS $parameters and UNION annotation - harness/mutation.py: deterministic mutants of the reference; hidden tests must fail - G0022 revalidated with harness fixes: oracle 100, null 0 - results in docs/faz1-tasarim.md 11f Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
21
CLAUDE.md
21
CLAUDE.md
@@ -80,7 +80,8 @@ python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N
|
||||
python3 -m harness.cli rescore runs/<run_dir>
|
||||
python3 -m harness.cli cleanup-list <PREFIX> runs/<dir> && python3 -m harness.cli teardown runs/<dir>
|
||||
python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000
|
||||
python3 -m harness.pilot 2
|
||||
python3 -m harness.pilot 2 # pilot list; `python3 -m harness.pilot 22 hard 1400` hard batch
|
||||
python3 -m harness.mutation G0002 --pool eval --run-base 5000 [--keep]
|
||||
python3 -c "from harness.ledger import spent; print(spent())"
|
||||
```
|
||||
|
||||
@@ -89,18 +90,18 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
||||
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support;
|
||||
task generator; ledger and budget guard.
|
||||
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
|
||||
- Pilot (G0002–G0021) done: 16/20 accepted, 0.16 USD per accepted task (`runs/gen/pilot.json`,
|
||||
analysis in `docs/faz1-tasarim.md` 11e). Generator improved after the pilot: static checks
|
||||
(name length, seed type, reserved words, contract/test classes), local abaplint parser check before SAP,
|
||||
max_tokens 80k, G2 detail in the repair feedback. Runner: G2 works for CDS views with parameters.
|
||||
The improved generator is not yet tested with new tasks.
|
||||
- G0020 (rejected) passes with oracle 100 after one fix (remove parameter `default`). G0013 (accepted) has a test
|
||||
class in its contract: fix in review.
|
||||
- Pilot (G0002–G0021): 16/20 accepted, 0.16 USD per accepted task (`docs/faz1-tasarim.md` 11e).
|
||||
G0020 and G0013 fixed in review and revalidated → 18/20.
|
||||
- Hard batch (G0022–G0026, 2026-10-03): 5/5 accepted (G0022 after harness fixes), 0.097 USD per accepted
|
||||
task, first-attempt 1/5 (`docs/faz1-tasarim.md` 11f).
|
||||
- Generator: static checks, local abaplint parser check, dependency order of seed/reference, budget floor,
|
||||
max_tokens 80k. Runner: G2 for CDS with parameters, G6 ignores unknown standard superclasses.
|
||||
- Step D: `harness/mutation.py` (deterministic mutants of the reference, hidden tests must fail).
|
||||
|
||||
## 7. Next steps
|
||||
|
||||
1. Test the improved generator with a small batch (~5 tasks, ~0.7 USD); compare repairs per task with the pilot (2.75 calls).
|
||||
2. Step D: mutation check (hidden tests must fail on a broken reference).
|
||||
1. Step D: run the mutation check on all accepted tasks; integrate it into the generator.
|
||||
2. Use killed mutants as faulty references for own-test scoring (component (a)).
|
||||
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
|
||||
checklist, Kral spot-checks ~10 flagged tasks.
|
||||
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
|
||||
|
||||
Reference in New Issue
Block a user