From f5611cbf56a0874960bbfec880744f45c8a81b96 Mon Sep 17 00:00:00 2001 From: Kral Date: Tue, 6 Oct 2026 05:56:31 +0200 Subject: [PATCH] data analysis doc: corrected two counts --- docs/data-analysis-stage2.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/data-analysis-stage2.md b/docs/data-analysis-stage2.md index 990fe5f..a2c71b0 100644 --- a/docs/data-analysis-stage2.md +++ b/docs/data-analysis-stage2.md @@ -33,7 +33,7 @@ run first for CLAS and FUNC to get more of these. Other frequent errors (top lis The teacher repairs after the first error in almost every case, one call later. It writes after a median of 3 reads. The longest search run of the teacher is 7, of Qwen 61. These are exactly the two behaviors the SFT must teach; both are in the data (repair share 60 to 65 % of accepted trajectories). Note: the Qwen figure includes failed runs, the teacher figure only accepted ones (a teacher run that never repaired is not accepted), so part of the gap is selection. -For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (30 runs), not analysed here. +For a fair view the rejected teacher runs would be added; they are in `runs/traj/` (29 runs), not analysed here. ## 3. Near duplicates @@ -45,7 +45,7 @@ between the two trajectories of one task. The data is not repetitive at this lev By kind: {'CLAS': 2, 'DDLS': 6, 'PROG': 3, 'TABL': 2} (CDS 6 of 13). The last call before the empty turn was a read in 10 of 13 cases ({'sap_push_source': 1, 'sap_syntax_check': 1, 'sap_pull_source': 4, 'sap_object_structure': 1, 'sap_search_object': 1, 'sap_sql_query': 4, 'sap_check_object': 1}), with small results (50 to 6000 characters) and contexts from 13k to 63k tokens: it is not a big tool result and not the context size. In every case the turn used the whole output limit (32000 tokens) on reasoning with no content and no tool call; in 7 of 13 runs two or three turns in a row did that (the retry with the same prompt and temperature reproduces it). Over all 1997 turns: p50 1109 tokens, p90 7862, -the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 12 longer than 20000, 22 longer than 16000. +the legitimate long turns (accepted runs, content produced) reach up to 30078; only 3 legitimate turns were longer than 24000, 10 longer than 20000, 22 longer than 16000. Fix: `docs/empty-response.md`. ## 5. What it means for the data