4.4 KiB
Stage 2 data analysis (accepted DeepSeek trajectories), 2026-10-06
Script train/analyze.py, numbers in runs/analysis/analysis.json. Base: 90 accepted trajectories of 119 runs (acceptance 76 %).
1. Repair taxonomy: are the three step-3 errors covered?
133 failed writes occurred inside accepted trajectories; each of the three common errors appears and is repaired:
| error (roadmap step 3) | failed writes | trajectories | repaired in the same trajectory | example |
|---|---|---|---|---|
TYPE c LENGTH n in a method signature |
10 | 3 | 3 | Unable to interpret "4". Possible causes of error include incorrect spellings or comma errors. |
| reserved word as a field or parameter name | 18 | 16 | 16 | Field "VALUE" is unknown. |
| name longer than 30 characters | 25 | 23 | 23 | The name "SORTS_BY_TOTAL_DESC_THEN_VARIETY" is longer than the allowed 30 characters. |
Reading: the long names (23 trajectories) and the reserved words (16, by a text match: the label is generous, it also counts "Field VALUE is unknown") are well covered.
TYPE c LENGTH in a signature is thin: 3 trajectories (10 failed writes, all repaired). The SAP message is only "save operation failed" or, with the proxy hint,
Statement does not exist ... METHODS. This is the error that the proxy syntaxCheck exists for; after the reset the error-targeted slots ("named-type") should be
run first for CLAS and FUNC to get more of these. Other frequent errors (top list in the json): the testclasses include is not reachable with sap_push_element (9 times),
"save operation failed" without detail (30 with the hint), "Field X is unknown", "Test method can be defined only in test classes".
2. Behaviors that Qwen lacks (series A, all 20 runs) versus the teacher (accepted trajectories)
| behavior | DeepSeek (accepted) | Qwen series A |
|---|---|---|
| runs with at least one failed write | 60 of 90 | 14 of 20 |
| repair after the first error (a different source that is written successfully) | 59 of 60 (98 %) | 6 of 14 (43 %) |
| median calls from the first error to the repair | 1 | 2.0 |
| searches/reads before the first write (median / p90) | 3.0 / 19 | 4.0 / 15 |
longest series of sap_search_object calls |
7 | 61 |
| runs with the same source pushed again | 2 of 90 | 11 of 20 |
| runs without any write | 4 | 4 |
The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61.
These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories).
Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection.
For a fair view the rejected teacher runs would be added; they are in runs/traj/ (29 runs), not analysed here.
3. Near duplicates
Pairwise comparison of the code that each accepted trajectory wrote (5-word shingles, prefix removed) and of the final reports: no pair above 0.8 (code) or 0.85 (report), also none between the two trajectories of one task. The data is not repetitive at this level.
4. empty_response (13 of 119 runs, 11 %)
By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and
contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call;
in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862,
the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 10 longer than 20000, 22 longer than 16000.
Fix: docs/empty-response.md.
5. What it means for the data
- Keep the repair behavior and the "write soon" behavior: they are present. Weight the error-targeted slots (named-type) up after the reset.
- Do not rely on
TYPE c LENGTHrepairs: only 3 examples. - The DDLS share of empty runs (6 of 13) explains much of the CDS rejection.