Restart plan step 0b: rerun of the four empirical filter runs (G0183 G0180 G0143 G0185) without touching the eval set; empirical --rerun
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -23,6 +23,17 @@ Ollama reset on 12 October. Until then only no-cloud work. This is the plan for
|
||||
or exception task, so the trained model could not be measured on them. Each slot is checked against the training pool. Review: `docs/eval-spotcheck-new-kinds.md`
|
||||
(Kral checks about 10 tasks). Run it before the training tasks so that the training tasks can be checked against the finished eval tasks.
|
||||
|
||||
## 1c. Step 0b: rerun of four empirical filter runs (Kral + Opus 2026-10-06)
|
||||
|
||||
Four eval candidates saw or read leftover objects of other runs in the first empirical filter (`docs/foreign-objects-report.md`): **G0183, G0180, G0143, G0185** (DeepSeek, scores 80, 85, 85, 85).
|
||||
After step 0 (eval generation for the new kinds), with the fixed proxy:
|
||||
|
||||
`python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 450000 --rerun G0183 G0180 G0143 G0185`
|
||||
|
||||
This writes only to `runs/emp_rerun/` (4 runs, about 1 ledger) and prints old and new filter decision per task. `empirical.json` and the eval set are **not** touched. Decision rules (step F):
|
||||
easy candidate = score 95 or more and at most 20 tool calls; flag = score under 50 (a strong model fails although the reference passes); else normal. **If a decision changes for one of the four, report it to Kral and
|
||||
Opus first; the eval set is changed only after their answer.** If nothing changes, say so in `docs/eval-inceleme.md` and leave the old results.
|
||||
|
||||
## 2. Order of work (what the controller does by itself)
|
||||
|
||||
Only kinds below their target share are generated and run (`harness/mix.py`, `below_target`); the kind with the biggest
|
||||
|
||||
Reference in New Issue
Block a user