21 lines
1.6 KiB
Markdown
21 lines
1.6 KiB
Markdown
# Own-test mutation score of the accepted trajectories (item D, metadata only)
|
|
|
|
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
|
|
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
|
|
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
|
|
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
|
|
|
|
## Result
|
|
90 accepted trajectories looked at: {'scored': 63, 'not_supported': 5, 'no_own_tests': 22}.
|
|
- Scored: 63; the model's tests also passed on the correct reference: 53 (the others are not reliable: a test that fails on the correct solution kills every mutant).
|
|
- Mean score over the reliable ones: 0.99; median 1.00; distribution {'1.0': 50, '>=0.75': 3} (n = 53).
|
|
- By kind (mean, n): {CLAS: 1.00 (45), DDLS: 0.90 (5), FUNC: 1.00 (3)}
|
|
- `no_own_tests`: 22 trajectories were accepted without any own test (at most 85 points); `no_mutants`: 0; `not_supported` (PROG: tests are inside the program): 5.
|
|
|
|
## Weakest (score under 0.75, tests pass on the reference)
|
|
|
|
|
|
## Use
|
|
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
|
|
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).
|