B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
55
docs/data-analysis-stage2.md
Normal file
55
docs/data-analysis-stage2.md
Normal file
@@ -0,0 +1,55 @@
|
||||
# Stage 2 data analysis (accepted DeepSeek trajectories), 2026-10-06
|
||||
|
||||
Script `train/analyze.py`, numbers in `runs/analysis/analysis.json`. Base: 90 accepted trajectories of 119 runs (acceptance 76 %).
|
||||
|
||||
## 1. Repair taxonomy: are the three step-3 errors covered?
|
||||
|
||||
133 failed writes occurred inside accepted trajectories; each of the three common errors appears and is repaired:
|
||||
|
||||
| error (roadmap step 3) | failed writes | trajectories | repaired in the same trajectory | example |
|
||||
|---|---|---|---|---|
|
||||
| `TYPE c LENGTH n` in a method signature | 10 | 3 | 3 | Unable to interpret "4". Possible causes of error include incorrect spellings or comma errors. |
|
||||
| reserved word as a field or parameter name | 18 | 16 | 16 | Field "VALUE" is unknown. |
|
||||
| name longer than 30 characters | 25 | 23 | 23 | The name "SORTS_BY_TOTAL_DESC_THEN_VARIETY" is longer than the allowed 30 characters. |
|
||||
|
||||
Reading: the long names (23 trajectories) and the reserved words (16, by a text match: the label is generous, it also counts "Field VALUE is unknown") are well covered.
|
||||
**`TYPE c LENGTH` in a signature is thin: 3 trajectories** (10 failed writes, all repaired). The SAP message is only "save operation failed" or, with the proxy hint,
|
||||
`Statement does not exist ... METHODS`. This is the error that the proxy syntaxCheck exists for; after the reset the error-targeted slots ("named-type") should be
|
||||
run first for CLAS and FUNC to get more of these. Other frequent errors (top list in the json): the testclasses include is not reachable with `sap_push_element` (9 times),
|
||||
"save operation failed" without detail (30 with the hint), "Field X is unknown", "Test method can be defined only in test classes".
|
||||
|
||||
## 2. Behaviors that Qwen lacks (series A, all 20 runs) versus the teacher (accepted trajectories)
|
||||
|
||||
| behavior | DeepSeek (accepted) | Qwen series A |
|
||||
|---|---|---|
|
||||
| runs with at least one failed write | 60 of 90 | 14 of 20 |
|
||||
| **repair after the first error** (a different source that is written successfully) | **59 of 60 (98 %)** | **6 of 14 (43 %)** |
|
||||
| median calls from the first error to the repair | 1 | 2.0 |
|
||||
| searches/reads before the first write (median / p90) | 3.0 / 19 | 4.0 / 15 |
|
||||
| longest series of `sap_search_object` calls | 7 | **61** |
|
||||
| runs with the same source pushed again | 2 of 90 | **11 of 20** |
|
||||
| runs without any write | 4 | 4 |
|
||||
|
||||
The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61.
|
||||
These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories).
|
||||
Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection.
|
||||
For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (30 runs), not analysed here.
|
||||
|
||||
## 3. Near duplicates
|
||||
|
||||
Pairwise comparison of the code that each accepted trajectory wrote (5-word shingles, prefix removed) and of the final reports: **no pair above 0.8 (code) or 0.85 (report)**, also none
|
||||
between the two trajectories of one task. The data is not repetitive at this level.
|
||||
|
||||
## 4. empty_response (13 of 119 runs, 11 %)
|
||||
|
||||
By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and
|
||||
contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call;
|
||||
in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862,
|
||||
the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 12 longer than 20000, 22 longer than 16000.
|
||||
Fix: `docs/empty-response.md`.
|
||||
|
||||
## 5. What it means for the data
|
||||
|
||||
- Keep the repair behavior and the "write soon" behavior: they are present. Weight the error-targeted slots (named-type) up after the reset.
|
||||
- Do not rely on `TYPE c LENGTH` repairs: only 3 examples.
|
||||
- The DDLS share of empty runs (6 of 13) explains much of the CDS rejection.
|
||||
23
docs/empty-response.md
Normal file
23
docs/empty-response.md
Normal file
@@ -0,0 +1,23 @@
|
||||
# empty_response fix (2026-10-06)
|
||||
|
||||
**Problem.** 13 of 119 DeepSeek runs (11 %) ended with `empty_response`: one turn used the whole output limit (32000 tokens) for reasoning,
|
||||
with no content and no tool call, and the retry (same prompt, same temperature, up to two) did it again in 7 of 13 runs. Each such run costs up to
|
||||
3 x 32000 output tokens and is lost (DDLS 6, PROG 3, CLAS 2, TABL 2). Analysis: `docs/data-analysis-stage2.md`, section 4.
|
||||
|
||||
**What the data says.** The empty turn comes after a small read result (50 to 6000 characters), in contexts of 13k to 63k tokens, not after a large result. Turns
|
||||
over all runs: p50 1.1k, p90 7.9k tokens; legitimate long turns (large test classes) reach 30k; only 3 of 1997 turns longer than 24k were legitimate.
|
||||
|
||||
**Implemented (harness side only, `harness/agents.py`, `harness/trajectories.py`)**
|
||||
1. Output cap per turn **24000** (was 32000; saves a quarter of every runaway turn, costs 3 of 1997 legitimate turns).
|
||||
2. **One retry at temperature 0.8** (was two retries at 0.2): another sample instead of the same runaway.
|
||||
3. **Stream guard** (off until verified): the request is streamed; when a turn has produced only reasoning for `STREAM_GUARD` tokens (suggested 9000) and no content and
|
||||
no tool call, the stream is cut and the turn counts as empty (retry as in 2). Saves most of the cost of a runaway turn. Tested with a fake streaming server (text, tool call,
|
||||
runaway: cut after 1500 estimated tokens in 0.01 s). **Not tested on the cloud model**: it needs the Ollama cloud stream to carry the reasoning in `delta.reasoning`
|
||||
(or `reasoning_content` / `thinking`) and tool calls as streamed deltas; if the stream cannot be parsed the agent falls back to the normal request. Usage of a cut turn is estimated
|
||||
(characters / 3.2) and marked `estimated` in the record.
|
||||
4. Not implemented: "one object per write call" rule and splitting large test classes. Both change what the model sees (system prompt or task) and so the training distribution and
|
||||
the comparison with the baseline (same system prompt). Option if 1 to 3 are not enough: a hint in the task spec of CDS tasks only.
|
||||
|
||||
**Verify after the reset (5 runs, DDLS and PROG first because they had most empty runs).** Run with `STREAM_GUARD=9000 python3 -m harness.pipeline` or set the constant:
|
||||
check that `empty_response` stays below 5 % over 40 runs, that no accepted run was cut wrongly (`cut_by_stream_guard` in `metadata.turn_usage` followed by a good turn is fine),
|
||||
and that the cost per run does not rise. If the stream breaks (tool calls missing), unset `STREAM_GUARD`: items 1 and 2 stay.
|
||||
Reference in New Issue
Block a user