Handover: no new model runs during the baseline; rerun queue
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -3,6 +3,12 @@
|
|||||||
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
|
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
|
||||||
CLAUDE.md and the docs when they are done, then empty this file.
|
CLAUDE.md and the docs when they are done, then empty this file.
|
||||||
|
|
||||||
|
## 0. Rule for the next session (Kral, 2026-10-03)
|
||||||
|
|
||||||
|
Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue
|
||||||
|
reruns for after the baseline in `runs/stage1/rerun_queue.txt` (one task id per line, with the reason).
|
||||||
|
The night chain still starts W7B and then the baseline as planned.
|
||||||
|
|
||||||
## 1. Running jobs
|
## 1. Running jobs
|
||||||
|
|
||||||
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
|
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
|
||||||
@@ -43,15 +49,19 @@ After the jobs:
|
|||||||
|
|
||||||
## 3. Pending follow-ups
|
## 3. Pending follow-ups
|
||||||
|
|
||||||
|
0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check
|
||||||
|
of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to
|
||||||
|
`runs/stage1/rerun_queue.txt`.
|
||||||
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
|
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
|
||||||
G0107). Old DeepSeek runs of FUNC tasks where G2 failed: re-score or rerun. G0107 is queued in W7B.
|
G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where
|
||||||
|
G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B.
|
||||||
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
|
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
|
||||||
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
|
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
|
||||||
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
|
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
|
||||||
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
|
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
|
||||||
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
|
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
|
||||||
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
|
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
|
||||||
response"). Not yet queued.
|
response"). Queued in `runs/stage1/rerun_queue.txt`.
|
||||||
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
|
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
|
||||||
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
|
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
|
||||||
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
|
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
|
||||||
|
|||||||
Reference in New Issue
Block a user