F: eval slots for INTF/TABL/STRU/MSAG/exception (+K), Kral spot-check sheet, step 0 in the restart plan; D: own-test mutation scoring (running)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -16,6 +16,13 @@ Ollama reset on 12 October. Until then only no-cloud work. This is the plan for
|
||||
Remove `runs/pipeline/STOP` and `STOPPED.txt` if they exist.
|
||||
4. Dashboard: `python3 -m harness.dashboard` (5-minute page, `runs/dashboard/index.html`).
|
||||
|
||||
## 1b. Step 0: eval tasks for the new kinds (before any training task)
|
||||
|
||||
`python3 -m harness.evalset plan-new` shows 32 slots (INTF 10, TABL 11, STRU 3, MSAG 4, exception 4; ids G0200 to G0231), then `python3 -m harness.evalset run-new`
|
||||
(about 5 ledger, about 2 usage), then `python3 -m harness.evalset run-new-k` (5 K variants). Reason: the eval set (130 accepted candidates) has no INTF, TABL, STRU, MSAG
|
||||
or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: `docs/eval-spotcheck-new-kinds.md`
|
||||
(Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.
|
||||
|
||||
## 2. Order of work (what the controller does by itself)
|
||||
|
||||
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest
|
||||
|
||||
Reference in New Issue
Block a user