# Stage 2 data analysis (accepted DeepSeek trajectories), 2026-10-06 Script `train/analyze.py`, numbers in `runs/analysis/analysis.json`. Base: 90 accepted trajectories of 119 runs (acceptance 76 %). ## 1. Repair taxonomy: are the three step-3 errors covered? 133 failed writes occurred inside accepted trajectories; each of the three common errors appears and is repaired: | error (roadmap step 3) | failed writes | trajectories | repaired in the same trajectory | example | |---|---|---|---|---| | `TYPE c LENGTH n` in a method signature | 10 | 3 | 3 | Unable to interpret "4". Possible causes of error include incorrect spellings or comma errors. | | reserved word as a field or parameter name | 18 | 16 | 16 | Field "VALUE" is unknown. | | name longer than 30 characters | 25 | 23 | 23 | The name "SORTS_BY_TOTAL_DESC_THEN_VARIETY" is longer than the allowed 30 characters. | Reading: the long names (23 trajectories) and the reserved words (16, by a text match: the label is generous, it also counts "Field VALUE is unknown") are well covered. **`TYPE c LENGTH` in a signature is thin: 3 trajectories** (10 failed writes, all repaired). The SAP message is only "save operation failed" or, with the proxy hint, `Statement does not exist ... METHODS`. This is the error that the proxy syntaxCheck exists for; after the reset the error-targeted slots ("named-type") should be run first for CLAS and FUNC to get more of these. Other frequent errors (top list in the json): the testclasses include is not reachable with `sap_push_element` (9 times), "save operation failed" without detail (30 with the hint), "Field X is unknown", "Test method can be defined only in test classes". ## 2. Behaviors that Qwen lacks (series A, all 20 runs) versus the teacher (accepted trajectories) | behavior | DeepSeek (accepted) | Qwen series A | |---|---|---| | runs with at least one failed write | 60 of 90 | 14 of 20 | | **repair after the first error** (a different source that is written successfully) | **59 of 60 (98 %)** | **6 of 14 (43 %)** | | median calls from the first error to the repair | 1 | 2.0 | | searches/reads before the first write (median / p90) | 3.0 / 19 | 4.0 / 15 | | longest series of `sap_search_object` calls | 7 | **61** | | runs with the same source pushed again | 2 of 90 | **11 of 20** | | runs without any write | 4 | 4 | The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61. These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories). Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection. For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (29 runs), not analysed here. ## 3. Near duplicates Pairwise comparison of the code that each accepted trajectory wrote (5-word shingles, prefix removed) and of the final reports: **no pair above 0.8 (code) or 0.85 (report)**, also none between the two trajectories of one task. The data is not repetitive at this level. ## 4. empty_response (13 of 119 runs, 11 %) By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call; in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862, the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 10 longer than 20000, 22 longer than 16000. Fix: `docs/empty-response.md`. ## 5. What it means for the data - Keep the repair behavior and the "write soon" behavior: they are present. Weight the error-targeted slots (named-type) up after the reset. - Do not rely on `TYPE c LENGTH` repairs: only 3 examples. - The DDLS share of empty runs (6 of 13) explains much of the CDS rejection.