Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 09:04:09 +02:00
parent a465c33e3f
commit 8ef85c2713
10 changed files with 402 additions and 41 deletions

View File

@@ -0,0 +1,29 @@
# Other runs' objects in tool results: which runs are affected (2026-10-06)
Cause: the model sometimes names its own helper or test class `ZCL_<run prefix>_...` (prefix inside the name). The teardown looked only for names that start with the prefix, so such
classes stayed in A4H and showed up in `sap_inactive_objects`, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, `harness/sweep.py`).
Scan: `train/foreign_scan.py` (read only; result `runs/analysis/foreign_scan.json`). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object.
| group | runs | with foreign names in a tool result | with a foreign read | tools |
|---|---|---|---|---|
| official baseline Qwen (MacBook, 11 tasks) | 11 | 0 | 0 | - |
| Devstral baseline (dropped) | 11 | 0 | 0 | - |
| older Qwen baselines (archive) | 8 | 0 | 0 | - |
| smoke | 1 | 0 | 0 | - |
| eval empirical filter (DeepSeek on eval candidates) | 178 | 25 | 4 | sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5 |
| eval generation validation (oracle, null, mutants) | 1344 | 0 | 0 | - |
| early pilot runs | 16 | 1 | 0 | sap_inactive_objects 1 |
| training trajectories (DeepSeek) | 117 | 67 | 15 | sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1 |
| series A (local Qwen) | 20 | 3 | 0 | sap_inactive_objects 2, sap_search_object 1 |
## Reading
- **Official baseline (Qwen, 11 tasks, 2026-10-04): not affected** (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers,
and none of the 11 baseline runs called `sap_inactive_objects` (checked: 0 calls). **The baseline numbers stay as they are. Nothing was changed.**
- **Eval generation (1344 validation runs: oracle, null, mutants): not affected** (they call no list or search tool).
- **Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one**
(G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there
was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs).
- **Early pilot runs:** 1 of 16 saw a foreign name in an inactive list.
- **Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one.** Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool
(`train/requeue_foreign.py`, rows in `runs/traj/summary_excluded.jsonl`) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder.
- **Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads.** The scores (3 of 20) are not affected by a read; the series ran before the fix.