STATE/handover: baseline restarted on the 11-task subset

This commit is contained in:
Kral
2026-10-03 22:36:30 +02:00
parent 689820af4d
commit 040b9900fd
2 changed files with 12 additions and 10 deletions

View File

@@ -1,4 +1,4 @@
# Handover notes (2026-10-03, 19:10)
# Handover notes (2026-10-03, 22:40; baseline shortened, see train/STATE.md)
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file.
@@ -30,7 +30,7 @@ Night chain (`train/night_chain.sh`):
`t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60):
`python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`.
If T01 failed: notification "T01 test failed: baseline not started", and it stops.
3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October
3. Notification "Baseline (11 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October
morning to noon.
After the jobs:
@@ -68,11 +68,11 @@ Done:
Waiting for jobs (do these in the next session):
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and
the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
- After "Baseline (11 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
Kral spot-check list (flagged tasks + one per category, ~10).
- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about
0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
- A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (25 tasks) ended"; no
- A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (11 tasks) ended"; no
second question needed): check `docker stats` and the HANA allocation limit first; choose the limit with the
data (about 32 GB is the SAP minimum for the trial image; too low gives an OOM kill); clean stop of `a4h`
(give HANA time), apply the limit, start, test `localhost:50000` and one MCP call. The Mac mini has 64 GB; swap