Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 07:07:44 +02:00
parent 040b9900fd
commit 0401186a6c
6 changed files with 79 additions and 8 deletions

View File

@@ -36,6 +36,11 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- The earlier night chain (25 tasks) was stopped with its first T01 run (aborted at 33 tool calls after 2 h;
A4H objects deleted; directory `runs/stage1/baseline/_aborted_20100_T01`).
- 2026-10-04: A4H memory: Docker VM 36 GB (was 64 GB = all RAM), container 32 GB; swap went from 15 GB to 1.4 GB.
Baseline switched to thinking off (`enable_thinking: false` per request); thinking-on results kept in
`runs/stage1/baseline_thinking_on.json`. Loop guard (3 identical pushes) and per-run activation error records
added (`harness/agents.py`, `harness/proxy.py`, `harness/runner.py`, `train/baseline.py`).
## Next
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.