Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -53,15 +53,18 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | on, `reasoning_effort` medium |
|
||||
| Thinking | **off (`enable_thinking: false`), fixed for baseline and after training.** Sent per request as `chat_template_kwargs` by `train/baseline.py` (the server default stays thinking on). Reason: with thinking on, all 3 baseline runs ended with empty responses at the thinking limit (`runs/stage1/baseline_thinking_on.json`) |
|
||||
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
||||
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
||||
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
||||
| Loop guard | `loop_guard` 3: the run ends when `sap_push_source` pushes the same source (object + md5) 3 times in a row; final report "Stopped: loop ...", `end_reason` "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it |
|
||||
| Per run record | `end_reason` (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), `activation_failures`, `activation_error_messages` (unique), also in `baseline.json` |
|
||||
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
|
||||
| Docker (A4H) | VM memory 36 GB (`MemoryMiB` 36864), container `--memory 32g --memory-swap 32g` (2026-10-04) |
|
||||
| Prompt cache of the server | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
||||
|
||||
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7) and the
|
||||
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) are not comparable.
|
||||
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7), the
|
||||
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thinking-on runs are not comparable.
|
||||
|
||||
## Baseline
|
||||
|
||||
|
||||
Reference in New Issue
Block a user