F: eval slots for INTF/TABL/STRU/MSAG/exception (+K), Kral spot-check sheet, step 0 in the restart plan; D: own-test mutation scoring (running)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 06:06:41 +02:00
parent 7f03849a86
commit 0c0e07fe96
4 changed files with 331 additions and 0 deletions

View File

@@ -16,6 +16,13 @@ Ollama reset on 12 October. Until then only no-cloud work. This is the plan for
Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist.
4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`).
## 1b. Step 0: eval tasks for the new kinds (before any training task)
`python3 -m harness.evalset plan-new` shows 32 slots (INTF 10, TABL 11, STRU 3, MSAG 4, exception 4; ids G0200 to G0231), then `python3 -m harness.evalset run-new`
(about 5 ledger, about 2 usage), then `python3 -m harness.evalset run-new-k` (5 K variants). Reason: the eval set (130 accepted candidates) has no INTF, TABL, STRU, MSAG
or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: `docs/eval-spotcheck-new-kinds.md`
(Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.
## 2. Order of work (what the controller does by itself)
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest