1.6 KiB
Own-test mutation score of the accepted trajectories (item D, metadata only)
Module harness/owntests.py, results runs/traj/<run>/own_test_mutation.json (not in git), report by train/own_test_report.py.
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (faulty/, mutants that the hidden tests kill).
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
Metadata only: the acceptance filter is not changed. The builder takes the score through --hook hooks_example:own_test_weight (field own_test_mutation).
Result
90 accepted trajectories looked at: {'scored': 63, 'not_supported': 5, 'no_own_tests': 22}.
- Scored: 63; the model's tests also passed on the correct reference: 53 (the others are not reliable: a test that fails on the correct solution kills every mutant).
- Mean score over the reliable ones: 0.99; median 1.00; distribution {'1.0': 50, '>=0.75': 3} (n = 53).
- By kind (mean, n): {CLAS: 1.00 (45), DDLS: 0.90 (5), FUNC: 1.00 (3)}
no_own_tests: 22 trajectories were accepted without any own test (at most 85 points);no_mutants: 0;not_supported(PROG: tests are inside the program): 5.
Weakest (score under 0.75, tests pass on the reference)
Use
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with tests_pass_on_reference false;
prefer trajectories without own tests last. The hook example shows where the weight goes (train/hooks_example.py).