Files
abap-llm/docs/empty-response.md

2.6 KiB

empty_response fix (2026-10-06)

Problem. 13 of 119 DeepSeek runs (11 %) ended with empty_response: one turn used the whole output limit (32000 tokens) for reasoning, with no content and no tool call, and the retry (same prompt, same temperature, up to two) did it again in 7 of 13 runs. Each such run costs up to 3 x 32000 output tokens and is lost (DDLS 6, PROG 3, CLAS 2, TABL 2). Analysis: docs/data-analysis-stage2.md, section 4.

What the data says. The empty turn comes after a small read result (50 to 6000 characters), in contexts of 13k to 63k tokens, not after a large result. Turns over all runs: p50 1.1k, p90 7.9k tokens; legitimate long turns (large test classes) reach 30k; only 3 of 1997 turns longer than 24k were legitimate.

Implemented (harness side only, harness/agents.py, harness/trajectories.py)

  1. Output cap per turn 24000 (was 32000; saves a quarter of every runaway turn, costs 3 of 1997 legitimate turns).
  2. One retry at temperature 0.8 (was two retries at 0.2): another sample instead of the same runaway.
  3. Stream guard (off until verified): the request is streamed; when a turn has produced only reasoning for STREAM_GUARD tokens (suggested 9000) and no content and no tool call, the stream is cut and the turn counts as empty (retry as in 2). Saves most of the cost of a runaway turn. Tested with a fake streaming server (text, tool call, runaway: cut after 1500 estimated tokens in 0.01 s). Not tested on the cloud model: it needs the Ollama cloud stream to carry the reasoning in delta.reasoning (or reasoning_content / thinking) and tool calls as streamed deltas; if the stream cannot be parsed the agent falls back to the normal request. Usage of a cut turn is estimated (characters / 3.2) and marked estimated in the record.
  4. Not implemented: "one object per write call" rule and splitting large test classes. Both change what the model sees (system prompt or task) and so the training distribution and the comparison with the baseline (same system prompt). Option if 1 to 3 are not enough: a hint in the task spec of CDS tasks only.

Verify after the reset (5 runs, DDLS and PROG first because they had most empty runs). Run with STREAM_GUARD=9000 python3 -m harness.pipeline or set the constant: check that empty_response stays below 5 % over 40 runs, that no accepted run was cut wrongly (cut_by_stream_guard in metadata.turn_usage followed by a good turn is fine), and that the cost per run does not rise. If the stream breaks (tool calls missing), unset STREAM_GUARD: items 1 and 2 stay.