Files
abap-llm/docs/own-test-mutation.md

21 lines
1.6 KiB
Markdown

# Own-test mutation score of the accepted trajectories (item D, metadata only)
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
## Result
90 accepted trajectories looked at: {'scored': 63, 'not_supported': 5, 'no_own_tests': 22}.
- Scored: 63; the model's tests also passed on the correct reference: 53 (the others are not reliable: a test that fails on the correct solution kills every mutant).
- Mean score over the reliable ones: 0.99; median 1.00; distribution {'1.0': 50, '>=0.75': 3} (n = 53).
- By kind (mean, n): {CLAS: 1.00 (45), DDLS: 0.90 (5), FUNC: 1.00 (3)}
- `no_own_tests`: 22 trajectories were accepted without any own test (at most 85 points); `no_mutants`: 0; `not_supported` (PROG: tests are inside the program): 5.
## Weakest (score under 0.75, tests pass on the reference)
## Use
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).