B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 05:56:23 +02:00
parent 1ddb9a7c65
commit 542fc8ec31
6 changed files with 406 additions and 9 deletions

View File

@@ -0,0 +1,55 @@
# Stage 2 data analysis (accepted DeepSeek trajectories), 2026-10-06
Script `train/analyze.py`, numbers in `runs/analysis/analysis.json`. Base: 90 accepted trajectories of 119 runs (acceptance 76 %).
## 1. Repair taxonomy: are the three step-3 errors covered?
133 failed writes occurred inside accepted trajectories; each of the three common errors appears and is repaired:
| error (roadmap step 3) | failed writes | trajectories | repaired in the same trajectory | example |
|---|---|---|---|---|
| `TYPE c LENGTH n` in a method signature | 10 | 3 | 3 | Unable to interpret "4". Possible causes of error include incorrect spellings or comma errors. |
| reserved word as a field or parameter name | 18 | 16 | 16 | Field "VALUE" is unknown. |
| name longer than 30 characters | 25 | 23 | 23 | The name "SORTS_BY_TOTAL_DESC_THEN_VARIETY" is longer than the allowed 30 characters. |
Reading: the long names (23 trajectories) and the reserved words (16, by a text match: the label is generous, it also counts "Field VALUE is unknown") are well covered.
**`TYPE c LENGTH` in a signature is thin: 3 trajectories** (10 failed writes, all repaired). The SAP message is only "save operation failed" or, with the proxy hint,
`Statement does not exist ... METHODS`. This is the error that the proxy syntaxCheck exists for; after the reset the error-targeted slots ("named-type") should be
run first for CLAS and FUNC to get more of these. Other frequent errors (top list in the json): the testclasses include is not reachable with `sap_push_element` (9 times),
"save operation failed" without detail (30 with the hint), "Field X is unknown", "Test method can be defined only in test classes".
## 2. Behaviors that Qwen lacks (series A, all 20 runs) versus the teacher (accepted trajectories)
| behavior | DeepSeek (accepted) | Qwen series A |
|---|---|---|
| runs with at least one failed write | 60 of 90 | 14 of 20 |
| **repair after the first error** (a different source that is written successfully) | **59 of 60 (98 %)** | **6 of 14 (43 %)** |
| median calls from the first error to the repair | 1 | 2.0 |
| searches/reads before the first write (median / p90) | 3.0 / 19 | 4.0 / 15 |
| longest series of `sap_search_object` calls | 7 | **61** |
| runs with the same source pushed again | 2 of 90 | **11 of 20** |
| runs without any write | 4 | 4 |
The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61.
These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories).
Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection.
For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (30 runs), not analysed here.
## 3. Near duplicates
Pairwise comparison of the code that each accepted trajectory wrote (5-word shingles, prefix removed) and of the final reports: **no pair above 0.8 (code) or 0.85 (report)**, also none
between the two trajectories of one task. The data is not repetitive at this level.
## 4. empty_response (13 of 119 runs, 11 %)
By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and
contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call;
in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862,
the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 12 longer than 20000, 22 longer than 16000.
Fix: `docs/empty-response.md`.
## 5. What it means for the data
- Keep the repair behavior and the "write soon" behavior: they are present. Weight the error-targeted slots (named-type) up after the reset.
- Do not rely on `TYPE c LENGTH` repairs: only 3 examples.
- The DDLS share of empty runs (6 of 13) explains much of the CDS rejection.