Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -53,15 +53,18 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | on, `reasoning_effort` medium |
|
||||
| Thinking | **off (`enable_thinking: false`), fixed for baseline and after training.** Sent per request as `chat_template_kwargs` by `train/baseline.py` (the server default stays thinking on). Reason: with thinking on, all 3 baseline runs ended with empty responses at the thinking limit (`runs/stage1/baseline_thinking_on.json`) |
|
||||
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
||||
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
||||
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
||||
| Loop guard | `loop_guard` 3: the run ends when `sap_push_source` pushes the same source (object + md5) 3 times in a row; final report "Stopped: loop ...", `end_reason` "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it |
|
||||
| Per run record | `end_reason` (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), `activation_failures`, `activation_error_messages` (unique), also in `baseline.json` |
|
||||
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
|
||||
| Docker (A4H) | VM memory 36 GB (`MemoryMiB` 36864), container `--memory 32g --memory-swap 32g` (2026-10-04) |
|
||||
| Prompt cache of the server | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
||||
|
||||
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7) and the
|
||||
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) are not comparable.
|
||||
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7), the
|
||||
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thinking-on runs are not comparable.
|
||||
|
||||
## Baseline
|
||||
|
||||
|
||||
@@ -36,6 +36,11 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
- The earlier night chain (25 tasks) was stopped with its first T01 run (aborted at 33 tool calls after 2 h;
|
||||
A4H objects deleted; directory `runs/stage1/baseline/_aborted_20100_T01`).
|
||||
|
||||
- 2026-10-04: A4H memory: Docker VM 36 GB (was 64 GB = all RAM), container 32 GB; swap went from 15 GB to 1.4 GB.
|
||||
Baseline switched to thinking off (`enable_thinking: false` per request); thinking-on results kept in
|
||||
`runs/stage1/baseline_thinking_on.json`. Loop guard (3 identical pushes) and per-run activation error records
|
||||
added (`harness/agents.py`, `harness/proxy.py`, `harness/runner.py`, `train/baseline.py`).
|
||||
|
||||
## Next
|
||||
|
||||
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
|
||||
|
||||
@@ -17,6 +17,7 @@ from harness.runner import Runner # noqa: E402
|
||||
|
||||
MODEL = os.path.expanduser("~/models/Qwen3.8-27B-4bit")
|
||||
BASE_URL = "http://127.0.0.1:8080/v1"
|
||||
LOOP_GUARD = 3 # end the run after 3 identical pushes in a row (end_reason "loop"); same for after training
|
||||
MAX_TOKENS = 16384 # per turn; sent in each request (the server limit stays 32768). Same for before/after.
|
||||
|
||||
|
||||
@@ -31,15 +32,16 @@ def main():
|
||||
out_path = os.path.join(ROOT, "runs", "stage1", f"{a.label}.json")
|
||||
os.makedirs(os.path.dirname(out_path), exist_ok=True)
|
||||
res = json.load(open(out_path)) if os.path.exists(out_path) else {"model": MODEL, "tasks": {}}
|
||||
res["settings"] = {"max_tokens": MAX_TOKENS, "reasoning_effort": "medium", "temperature": 0.2, "tool_call_budget": 60,
|
||||
"subset": "train/subset.json", "date": "2026-10-03"}
|
||||
res["settings"] = {"max_tokens": MAX_TOKENS, "enable_thinking": False, "temperature": 0.2, "tool_call_budget": 60, "loop_guard": LOOP_GUARD,
|
||||
"subset": "train/subset.json", "date": "2026-10-04"}
|
||||
runs_root = os.path.join(ROOT, "runs", "stage1", a.label)
|
||||
for i, t in enumerate(subset):
|
||||
tid = t["id"]
|
||||
if (a.only and tid not in a.only) or tid in res["tasks"]:
|
||||
continue
|
||||
pool = os.path.join(ROOT, "tasks") if tid.startswith("T") else os.path.join(ROOT, "tasks_gen", "eval")
|
||||
agent = LlmAgent(MODEL, BASE_URL, max_tokens=MAX_TOKENS)
|
||||
agent = LlmAgent(MODEL, BASE_URL, max_tokens=MAX_TOKENS, chat_template_kwargs={"enable_thinking": False},
|
||||
loop_guard=LOOP_GUARD)
|
||||
try:
|
||||
rep, run_dir = Runner(pool, runs_root).run(tid, agent, a.run_base + i)
|
||||
except Exception as e: # noqa: BLE001 one broken run must not stop the series
|
||||
@@ -52,6 +54,8 @@ def main():
|
||||
"score": (rep.get("score") or {}).get("total"), "parts": rep.get("score"),
|
||||
"gates": rep.get("gates"), "hidden": f"{h.get('passed')}/{h.get('total')}",
|
||||
"tool_calls": rep.get("tool_calls"), "seconds": rep.get("seconds"),
|
||||
"end_reason": rep.get("end_reason"), "activation_failures": rep.get("activation_failures"),
|
||||
"activation_error_messages": rep.get("activation_error_messages"),
|
||||
"agent_seconds": rep.get("agent_seconds"), "final": (rep.get("final_report") or "")[:300],
|
||||
"run_dir": os.path.relpath(run_dir, ROOT)}
|
||||
json.dump(res, open(out_path, "w"), indent=1)
|
||||
|
||||
Reference in New Issue
Block a user