B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
23
docs/empty-response.md
Normal file
23
docs/empty-response.md
Normal file
@@ -0,0 +1,23 @@
|
||||
# empty_response fix (2026-10-06)
|
||||
|
||||
**Problem.** 13 of 119 DeepSeek runs (11 %) ended with `empty_response`: one turn used the whole output limit (32000 tokens) for reasoning,
|
||||
with no content and no tool call, and the retry (same prompt, same temperature, up to two) did it again in 7 of 13 runs. Each such run costs up to
|
||||
3 x 32000 output tokens and is lost (DDLS 6, PROG 3, CLAS 2, TABL 2). Analysis: `docs/data-analysis-stage2.md`, section 4.
|
||||
|
||||
**What the data says.** The empty turn comes after a small read result (50 to 6000 characters), in contexts of 13k to 63k tokens, not after a large result. Turns
|
||||
over all runs: p50 1.1k, p90 7.9k tokens; legitimate long turns (large test classes) reach 30k; only 3 of 1997 turns longer than 24k were legitimate.
|
||||
|
||||
**Implemented (harness side only, `harness/agents.py`, `harness/trajectories.py`)**
|
||||
1. Output cap per turn **24000** (was 32000; saves a quarter of every runaway turn, costs 3 of 1997 legitimate turns).
|
||||
2. **One retry at temperature 0.8** (was two retries at 0.2): another sample instead of the same runaway.
|
||||
3. **Stream guard** (off until verified): the request is streamed; when a turn has produced only reasoning for `STREAM_GUARD` tokens (suggested 9000) and no content and
|
||||
no tool call, the stream is cut and the turn counts as empty (retry as in 2). Saves most of the cost of a runaway turn. Tested with a fake streaming server (text, tool call,
|
||||
runaway: cut after 1500 estimated tokens in 0.01 s). **Not tested on the cloud model**: it needs the Ollama cloud stream to carry the reasoning in `delta.reasoning`
|
||||
(or `reasoning_content` / `thinking`) and tool calls as streamed deltas; if the stream cannot be parsed the agent falls back to the normal request. Usage of a cut turn is estimated
|
||||
(characters / 3.2) and marked `estimated` in the record.
|
||||
4. Not implemented: "one object per write call" rule and splitting large test classes. Both change what the model sees (system prompt or task) and so the training distribution and
|
||||
the comparison with the baseline (same system prompt). Option if 1 to 3 are not enough: a hint in the task spec of CDS tasks only.
|
||||
|
||||
**Verify after the reset (5 runs, DDLS and PROG first because they had most empty runs).** Run with `STREAM_GUARD=9000 python3 -m harness.pipeline` or set the constant:
|
||||
check that `empty_response` stays below 5 % over 40 runs, that no accepted run was cut wrongly (`cut_by_stream_guard` in `metadata.turn_usage` followed by a good turn is fine),
|
||||
and that the cost per run does not rise. If the stream breaks (tool calls missing), unset `STREAM_GUARD`: items 1 and 2 stay.
|
||||
Reference in New Issue
Block a user