Baseline restarted with loop guard, thinking off (run base 20400)

This commit is contained in:
Kral
2026-10-04 07:40:19 +02:00
parent 0401186a6c
commit a94c2a3b5d
2 changed files with 7 additions and 14 deletions

View File

@@ -28,18 +28,11 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
## Running (detached)
- MLX server PID 66098 (not restarted), log `runs/stage1/server.log`.
- Baseline on the 11-task subset (`train/subset.json`; the 25-task subset is `train/subset_v1_25.json`), started
2026-10-03 22:36 by `train/baseline_chain.sh` (PID 69101; python PID 69107), logs `runs/stage1/baseline_chain.log`
and `runs/stage1/baseline.log`. Settings: max_tokens 16384, reasoning medium, budget 60 (`train/README.md`).
macOS notification after 2 tasks ("ask for the time estimate") and at the end ("Baseline (11 tasks) ended").
- The earlier night chain (25 tasks) was stopped with its first T01 run (aborted at 33 tool calls after 2 h;
A4H objects deleted; directory `runs/stage1/baseline/_aborted_20100_T01`).
- 2026-10-04: A4H memory: Docker VM 36 GB (was 64 GB = all RAM), container 32 GB; swap went from 15 GB to 1.4 GB.
Baseline switched to thinking off (`enable_thinking: false` per request); thinking-on results kept in
`runs/stage1/baseline_thinking_on.json`. Loop guard (3 identical pushes) and per-run activation error records
added (`harness/agents.py`, `harness/proxy.py`, `harness/runner.py`, `train/baseline.py`).
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
2026-10-04 about 08:00 by `train/baseline_chain.sh` (run base 20400), logs `runs/stage1/baseline_chain.log` and
`runs/stage1/baseline.log`. Notification after 2 tasks and at the end ("Baseline (11 tasks) ended").
- The T01 trial runs without the guard were stopped (`_aborted_*` in `runs/stage1/baseline/`, A4H objects deleted).
## Next