From 040b9900fd7a383206a4f6dea08a1515b4e39b8c Mon Sep 17 00:00:00 2001 From: Kral Date: Sat, 3 Oct 2026 22:36:30 +0200 Subject: [PATCH] STATE/handover: baseline restarted on the 11-task subset --- docs/devir-notlari.md | 8 ++++---- train/STATE.md | 14 ++++++++------ 2 files changed, 12 insertions(+), 10 deletions(-) diff --git a/docs/devir-notlari.md b/docs/devir-notlari.md index aa9cf97..9ee2c02 100644 --- a/docs/devir-notlari.md +++ b/docs/devir-notlari.md @@ -1,4 +1,4 @@ -# Handover notes (2026-10-03, 19:10) +# Handover notes (2026-10-03, 22:40; baseline shortened, see train/STATE.md) Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into CLAUDE.md and the docs when they are done, then empty this file. @@ -30,7 +30,7 @@ Night chain (`train/night_chain.sh`): `t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60): `python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`. If T01 failed: notification "T01 test failed: baseline not started", and it stops. -3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October +3. Notification "Baseline (11 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October morning to noon. After the jobs: @@ -68,11 +68,11 @@ Done: Waiting for jobs (do these in the next session): - After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?). -- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the +- After "Baseline (11 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the Kral spot-check list (flagged tasks + one per category, ~10). - Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about 0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting. -- A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (25 tasks) ended"; no +- A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (11 tasks) ended"; no second question needed): check `docker stats` and the HANA allocation limit first; choose the limit with the data (about 32 GB is the SAP minimum for the trial image; too low gives an OOM kill); clean stop of `a4h` (give HANA time), apply the limit, start, test `localhost:50000` and one MCP call. The Mac mini has 64 GB; swap diff --git a/train/STATE.md b/train/STATE.md index c0215fc..3c1a534 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -1,6 +1,6 @@ # Stage 1 state -Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:25. +Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:40. ## Done @@ -28,11 +28,13 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U ## Running (detached) -- MLX server PID 59352, log `runs/stage1/server.log`. -- T01 test (budget 40, harness check), PID 59414, log `runs/stage1/baseline_t01.log`. -- Night chain `train/night_chain.sh` PID 60017, log `runs/stage1/night_chain.log`: after the DeepSeek reruns - and the T01 test, it starts the baseline on the 25 tasks (budget 60), log `runs/stage1/baseline.log`, - macOS notification at the end. Expected end: 4 October, morning to noon. +- MLX server PID 66098 (not restarted), log `runs/stage1/server.log`. +- Baseline on the 11-task subset (`train/subset.json`; the 25-task subset is `train/subset_v1_25.json`), started + 2026-10-03 22:36 by `train/baseline_chain.sh` (PID 69101; python PID 69107), logs `runs/stage1/baseline_chain.log` + and `runs/stage1/baseline.log`. Settings: max_tokens 16384, reasoning medium, budget 60 (`train/README.md`). + macOS notification after 2 tasks ("ask for the time estimate") and at the end ("Baseline (11 tasks) ended"). +- The earlier night chain (25 tasks) was stopped with its first T01 run (aborted at 33 tool calls after 2 h; + A4H objects deleted; directory `runs/stage1/baseline/_aborted_20100_T01`). ## Next