Handover: no new model runs during the baseline; rerun queue

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 18:43:25 +02:00
parent 9366ee9868
commit 81d8cf6f99

View File

@@ -3,6 +3,12 @@
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file. CLAUDE.md and the docs when they are done, then empty this file.
## 0. Rule for the next session (Kral, 2026-10-03)
Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue
reruns for after the baseline in `runs/stage1/rerun_queue.txt` (one task id per line, with the reason).
The night chain still starts W7B and then the baseline as planned.
## 1. Running jobs ## 1. Running jobs
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends. All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
@@ -43,15 +49,19 @@ After the jobs:
## 3. Pending follow-ups ## 3. Pending follow-ups
0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check
of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to
`runs/stage1/rerun_queue.txt`.
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes, 1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
G0107). Old DeepSeek runs of FUNC tasks where G2 failed: re-score or rerun. G0107 is queued in W7B. G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where
G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B.
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked 2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form. for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations 3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued (G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60. (W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model 4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
response"). Not yet queued. response"). Queued in `runs/stage1/rerun_queue.txt`.
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model 5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open). error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks