Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 22:23:04 +02:00
parent 0677d035da
commit 7b8ca01bde
5 changed files with 179 additions and 211 deletions

View File

@@ -14,7 +14,7 @@ Stage 2 (SFT on teacher trajectories) is not part of this task.
Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03): Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03):
SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`). SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`).
Real numbers from Step 1 (`train/data/report.md`): 370 records, 2,806,501 tokens (Qwen tokenizer; Real numbers from Step 1 (`train/data/report.md`): 370 records, 2,806,501 tokens (Qwen tokenizer;
the chars/4 estimate is 2.64M). After version dedup and splitting: train 374 documents / 2.57M tokens, the chars/4 estimate is 2.64M). After version dedup and splitting: train 374 documents / 2.53M tokens,
valid 20 documents / 135k tokens, 748 iterations for 2 epochs. The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line: valid 20 documents / 135k tokens, 748 iterations for 2 epochs. The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line:
`{"text", "objects", "types", "language_version", "package_path", `{"text", "objects", "types", "language_version", "package_path",
"grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`. "grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`.

View File

@@ -1,6 +1,6 @@
# Stage 1 state # Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:30. Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 23:00.
## Done ## Done
@@ -15,14 +15,16 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`, - Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`). run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): `train/prepare.py`, report `train/data/report.md`. - Step 1 done (2026-10-03 23:00, Kral decisions applied, strict dedup rule): `train/prepare.py`, report
Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens. `train/data/report.md`. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept; Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with
tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed.
(classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82
over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %. pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834;
Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003). no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %.
**618 -> 748 iterations for 2 epochs.** `test.jsonl` = copy of valid. Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
`test.jsonl` = copy of valid.
## Running (detached) ## Running (detached)
@@ -34,6 +36,8 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
## Next ## Next
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json` 1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit. (`t01_test_budget40` holds the T01 test result). Commit.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline. 2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.

View File

@@ -3,24 +3,24 @@
Corpus: `/Users/erhankeseli/projects/abap-llm/corpus/corpus.jsonl` (SAP-samples/abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md). Corpus: `/Users/erhankeseli/projects/abap-llm/corpus/corpus.jsonl` (SAP-samples/abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md).
Tokenizer: `/Users/erhankeseli/models/Qwen3.8-27B-4bit`. Seed 20261003. max_seq_length 16384. Tokenizer: `/Users/erhankeseli/models/Qwen3.8-27B-4bit`. Seed 20261003. max_seq_length 16384.
## Version dedup (keep the newest; keep an older version only if it differs by more than 5 % of lines) ## Version dedup (strict rule; an older version is kept only if it differs by more than 5 % of lines)
- Object families: 269; with more than one version: 39. - Object families: 269; with more than one version: 39.
- Versions kept: 319. Versions removed: 16. - Versions kept: 323. Versions removed: 12.
- Records: 370 before, 351 after. - Records: 370 before, 353 after.
- Tokens before: 2806501 (chars/4 estimate 2636112). After dedup: 2699110. After splitting: 2701996. - Tokens before: 2806501 (chars/4 estimate 2636112). After dedup: 2658082. After splitting: 2660834.
- Info: a stricter rule (compare also with the other kept versions) would remove 8 more kept versions. - Rule: main (Cloud) and the newest v* (Standard ABAP) always kept; each older v* compared with the next newer v*; other branches (0 comparisons) compared with main.
## Splitting of documents over the limit ## Splitting of documents over the limit
- Documents split: 43; pieces: 86; hard line splits (a block was too big): 17; pieces still over the limit: 0. - Documents split: 41; pieces: 82; hard line splits (a block was too big): 15; pieces still over the limit: 0.
## Result ## Result
| | docs | tokens | | | docs | tokens |
|---|---|---| |---|---|---|
| train | 374 | 2567180 | | train | 374 | 2525568 |
| valid (= test file) | 20 | 134816 | | valid (= test file) | 20 | 135266 |
Split by family: 184 families. **Iterations for 2 epochs (batch 1): 748.** Split by family: 184 families. **Iterations for 2 epochs (batch 1): 748.**
@@ -28,54 +28,50 @@ Split by family: 184 families. **Iterations for 2 epochs (batch 1): 748.**
| | before dedup | final (kept pieces) | | | before dedup | final (kept pieces) |
|---|---|---| |---|---|---|
| DOC | 1301698 (46.4 %) | 1302728 (48.2 %) | | DOC | 1301698 (46.4 %) | 1302728 (49.0 %) |
| CLAS | 1378559 (49.1 %) | 1297878 (48.0 %) | | CLAS | 1378559 (49.1 %) | 1247458 (46.9 %) |
| other | 126244 (4.5 %) | 101390 (3.8 %) | | other | 126244 (4.5 %) | 110648 (4.2 %) |
## Token distribution per document (final) ## Token distribution per document (final)
| set | min | p50 | p90 | p99 | max | | set | min | p50 | p90 | p99 | max |
|---|---|---|---|---|---| |---|---|---|---|---|---|
| all | 49 | 4904 | 15608 | 16217 | 16273 | | all | 49 | 4797 | 15577 | 16217 | 16273 |
| train | 49 | 4904 | 15638 | 16177 | 16252 | | train | 49 | 4770 | 15590 | 16177 | 16252 |
| valid | 92 | 7589 | 15317 | 16273 | 16273 | | valid | 135 | 7589 | 15317 | 16273 | 16273 |
## By object type (final) ## By object type (final)
| types | docs | tokens | | types | docs | tokens |
|---|---|---| |---|---|---|
| DOC | 138 | 1302728 | | DOC | 138 | 1302728 |
| CLAS | 194 | 1297878 | | CLAS | 187 | 1247458 |
| PROG | 11 | 40071 | | PROG | 11 | 40071 |
| mixed(4) | 6 | 28405 | | mixed(4) | 6 | 28405 |
| DDLS | 21 | 11549 | | DDLS | 27 | 20614 |
| mixed(3) | 2 | 6667 | | mixed(3) | 2 | 6667 |
| DDLX | 6 | 5003 | | DDLX | 6 | 5003 |
| CLAS+DDLS | 1 | 4563 | | CLAS+DDLS | 1 | 4563 |
| FUGR | 2 | 2026 | | FUGR | 2 | 2026 |
| INTF | 7 | 1792 | | INTF | 7 | 1792 |
| BDEF | 4 | 1069 | | BDEF | 5 | 1262 |
| DCLS | 2 | 245 | | DCLS | 2 | 245 |
## Removed versions ## Removed versions
| object | version | newest | diff | tokens | | object | version | newest | diff | tokens |
|---|---|---|---|---| |---|---|---|---|---|
| ZBP_DEMO_ABAP_RAP_RO_U (CLAS) | v758 | main | 0.015 | 10696 | | ZCL_DEMO_ABAP_OBJECTS (CLAS) | v755 | v757 | 0.01 | 14944 |
| ZDEMO_ABAP_RAP_RO_M (BDEF) | v756 | main | 0.037 | 259 | | ZDEMO_ABAP_RAP_DRAFT_M (BDEF) | v756 | v757 | 0.033 | 302 |
| ZDEMO_ABAP_RAP_RO_U (BDEF) | v756 | main | 0.043 | 236 | | ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS) | v756 | v757 | 0.015 | 21263 |
| ZDEMO_ABAP_CDS_VE_ASSOC_E (DDLS) | v757 | main | 0.026 | 542 | | ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS) | v755 | v756 | 0.047 | 20431 |
| ZDEMO_ABAP_CDS_VE_JOINS (DDLS) | v757 | main | 0.009 | 1827 |
| ZDEMO_ABAP_CDS_VE_AGG_EXP (DDLS) | v757 | main | 0.019 | 804 |
| ZDEMO_ABAP_CDS_VE_SEL (DDLS) | v757 | main | 0.029 | 3273 |
| ZDEMO_ABAP_CDS_VE_ASSOC (DDLS) | v757 | main | 0.008 | 2451 |
| ZDEMO_ABAP_FLI_VE (DDLS) | v816 | main | 0.0 | 168 |
| ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS) | v816 | main | 0.009 | 27407 |
| ZCL_DEMO_ABAP_DISPLAY (CLAS) | v755 | v757 | 0.02 | 1677 | | ZCL_DEMO_ABAP_DISPLAY (CLAS) | v755 | v757 | 0.02 | 1677 |
| ZCL_DEMO_ABAP_UNIT_TEST (CLAS) | v758 | main | 0.034 | 19799 | | ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS) | v755 | v757 | 0.01 | 17309 |
| ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS) | v756 | v757 | 0.021 | 25883 |
| ZCL_DEMO_ABAP_SQL (CLAS) | v755 | v757 | 0.035 | 14949 |
| ZCL_DEMO_ABAP_CONSTRUCTOR_EXPR (CLAS) | v755 | v757 | 0.048 | 13893 |
| ZDEMO_ABAP_DYNPRO (PROG) | v757 | v816 | 0.001 | 8026 | | ZDEMO_ABAP_DYNPRO (PROG) | v757 | v816 | 0.001 | 8026 |
| ZBP_DEMO_ABAP_RAP_DRAFT_M (CLAS) | v757 | v758 | 0.005 | 2474 | | ZBP_DEMO_ABAP_RAP_DRAFT_M (CLAS) | v757 | v758 | 0.005 | 2474 |
| ZCL_DEMO_ABAP_NUMERIC_OP (CLAS) | v816 | main | 0.003 | 20484 |
| ZDEMO_ABAP_ALV (PROG) | v757 | v816 | 0.002 | 7268 | | ZDEMO_ABAP_ALV (PROG) | v757 | v816 | 0.002 | 7268 |
## Split documents ## Split documents
@@ -97,9 +93,8 @@ Split by family: 184 families. **Iterations for 2 epochs (batch 1): 748.**
- ['06_Dynamic_Programming__P07 (DOC)']: 23994 tokens -> 2 pieces [15836, 8221] - ['06_Dynamic_Programming__P07 (DOC)']: 23994 tokens -> 2 pieces [15836, 8221]
- ['28_Regular_Expressions (DOC)']: 17900 tokens -> 2 pieces [15031, 2930] - ['28_Regular_Expressions (DOC)']: 17900 tokens -> 2 pieces [15031, 2930]
- ['06_Dynamic_Programming__P06 (DOC)']: 16923 tokens -> 2 pieces [15842, 1144] - ['06_Dynamic_Programming__P06 (DOC)']: 16923 tokens -> 2 pieces [15842, 1144]
- ['ZCL_DEMO_ABAP_NUMERIC_OP__V816 (CLAS)']: 19769 tokens -> 2 pieces [15819, 4012]
- ['05_Constructor_Expressions__P01 (DOC)']: 17349 tokens -> 2 pieces [15147, 2265] - ['05_Constructor_Expressions__P01 (DOC)']: 17349 tokens -> 2 pieces [15147, 2265]
- ['ZCL_DEMO_ABAP_DTYPE_DOBJ__V755 (CLAS)']: 17237 tokens -> 2 pieces [15833, 1469]
- ['ZCL_DEMO_ABAP_DTYPE_DOBJ__V756 (CLAS)']: 18069 tokens -> 2 pieces [15842, 2292]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 17923 tokens -> 2 pieces [15329, 2672] - ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 17923 tokens -> 2 pieces [15329, 2672]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 16924 tokens -> 2 pieces [15749, 1223] - ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 16924 tokens -> 2 pieces [15749, 1223]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 16770 tokens -> 2 pieces [14210, 2608] - ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 16770 tokens -> 2 pieces [14210, 2608]
@@ -116,10 +111,9 @@ Split by family: 184 families. **Iterations for 2 epochs (batch 1): 748.**
- ['ZCL_DEMO_ABAP_INTERNAL_TABLES__V758 (CLAS)']: 17355 tokens -> 2 pieces [15832, 1584] - ['ZCL_DEMO_ABAP_INTERNAL_TABLES__V758 (CLAS)']: 17355 tokens -> 2 pieces [15832, 1584]
- ['ZCL_DEMO_ABAP_BUILTIN_FUNC (CLAS)']: 20446 tokens -> 2 pieces [15454, 5071] - ['ZCL_DEMO_ABAP_BUILTIN_FUNC (CLAS)']: 20446 tokens -> 2 pieces [15454, 5071]
- ['ZCL_DEMO_ABAP_STRING_PROC (CLAS)']: 20311 tokens -> 2 pieces [15590, 4797] - ['ZCL_DEMO_ABAP_STRING_PROC (CLAS)']: 20311 tokens -> 2 pieces [15590, 4797]
- ['ZCL_DEMO_ABAP_UNIT_TEST__V758 (CLAS)']: 16394 tokens -> 2 pieces [15301, 1187]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)']: 21483 tokens -> 2 pieces [15815, 5730] - ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)']: 21483 tokens -> 2 pieces [15815, 5730]
- ['ZCL_DEMO_ABAP_DATE_TIME (CLAS)']: 18634 tokens -> 2 pieces [15577, 3138] - ['ZCL_DEMO_ABAP_DATE_TIME (CLAS)']: 18634 tokens -> 2 pieces [15577, 3138]
- ['ZCL_DEMO_ABAP_XML_JSON (CLAS)']: 16825 tokens -> 2 pieces [15753, 1154] - ['ZCL_DEMO_ABAP_XML_JSON (CLAS)']: 16825 tokens -> 2 pieces [15753, 1154]
- ['ZCL_DEMO_ABAP_SQL (CLAS)']: 18283 tokens -> 2 pieces [12526, 5831] - ['ZCL_DEMO_ABAP_SQL (CLAS)']: 18283 tokens -> 2 pieces [12526, 5831]
- ['ZCL_DEMO_ABAP_INTERNAL_TABLES__V755 (CLAS)']: 17309 tokens -> 2 pieces [2124, 15283]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V756 (CLAS)']: 20979 tokens -> 2 pieces [15831, 5210]
- ['16_Data_Types_and_Objects__P02 (DOC)']: 17891 tokens -> 2 pieces [14421, 3541] - ['16_Data_Types_and_Objects__P02 (DOC)']: 17891 tokens -> 2 pieces [14421, 3541]

View File

@@ -5,84 +5,42 @@
"estimate_tokens": 2636112, "estimate_tokens": 2636112,
"records": 370, "records": 370,
"tokens_before_dedup": 2806501, "tokens_before_dedup": 2806501,
"tokens_after_dedup": 2699110, "tokens_after_dedup": 2658082,
"records_after_dedup": 351, "records_after_dedup": 353,
"tokens_after_split": 2701996, "tokens_after_split": 2660834,
"objects_with_versions": 39, "objects_with_versions": 39,
"families": 269, "families": 269,
"versions_kept": 319, "versions_kept": 323,
"versions_removed": 16, "versions_removed": 12,
"strict_rule_would_remove_more": 8, "other_branch_comparisons": 0,
"removed_versions": [ "removed_versions": [
{ {
"object": "ZBP_DEMO_ABAP_RAP_RO_U (CLAS)", "object": "ZCL_DEMO_ABAP_OBJECTS (CLAS)",
"version": "v758", "version": "v755",
"newest": "main", "newest": "v757",
"diff": 0.015, "diff": 0.01,
"tokens": 10696 "tokens": 14944
}, },
{ {
"object": "ZDEMO_ABAP_RAP_RO_M (BDEF)", "object": "ZDEMO_ABAP_RAP_DRAFT_M (BDEF)",
"version": "v756", "version": "v756",
"newest": "main", "newest": "v757",
"diff": 0.037, "diff": 0.033,
"tokens": 259 "tokens": 302
},
{
"object": "ZDEMO_ABAP_RAP_RO_U (BDEF)",
"version": "v756",
"newest": "main",
"diff": 0.043,
"tokens": 236
},
{
"object": "ZDEMO_ABAP_CDS_VE_ASSOC_E (DDLS)",
"version": "v757",
"newest": "main",
"diff": 0.026,
"tokens": 542
},
{
"object": "ZDEMO_ABAP_CDS_VE_JOINS (DDLS)",
"version": "v757",
"newest": "main",
"diff": 0.009,
"tokens": 1827
},
{
"object": "ZDEMO_ABAP_CDS_VE_AGG_EXP (DDLS)",
"version": "v757",
"newest": "main",
"diff": 0.019,
"tokens": 804
},
{
"object": "ZDEMO_ABAP_CDS_VE_SEL (DDLS)",
"version": "v757",
"newest": "main",
"diff": 0.029,
"tokens": 3273
},
{
"object": "ZDEMO_ABAP_CDS_VE_ASSOC (DDLS)",
"version": "v757",
"newest": "main",
"diff": 0.008,
"tokens": 2451
},
{
"object": "ZDEMO_ABAP_FLI_VE (DDLS)",
"version": "v816",
"newest": "main",
"diff": 0.0,
"tokens": 168
}, },
{ {
"object": "ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS)", "object": "ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS)",
"version": "v816", "version": "v756",
"newest": "main", "newest": "v757",
"diff": 0.009, "diff": 0.015,
"tokens": 27407 "tokens": 21263
},
{
"object": "ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS)",
"version": "v755",
"newest": "v756",
"diff": 0.047,
"tokens": 20431
}, },
{ {
"object": "ZCL_DEMO_ABAP_DISPLAY (CLAS)", "object": "ZCL_DEMO_ABAP_DISPLAY (CLAS)",
@@ -92,11 +50,32 @@
"tokens": 1677 "tokens": 1677
}, },
{ {
"object": "ZCL_DEMO_ABAP_UNIT_TEST (CLAS)", "object": "ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)",
"version": "v758", "version": "v755",
"newest": "main", "newest": "v757",
"diff": 0.034, "diff": 0.01,
"tokens": 19799 "tokens": 17309
},
{
"object": "ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)",
"version": "v756",
"newest": "v757",
"diff": 0.021,
"tokens": 25883
},
{
"object": "ZCL_DEMO_ABAP_SQL (CLAS)",
"version": "v755",
"newest": "v757",
"diff": 0.035,
"tokens": 14949
},
{
"object": "ZCL_DEMO_ABAP_CONSTRUCTOR_EXPR (CLAS)",
"version": "v755",
"newest": "v757",
"diff": 0.048,
"tokens": 13893
}, },
{ {
"object": "ZDEMO_ABAP_DYNPRO (PROG)", "object": "ZDEMO_ABAP_DYNPRO (PROG)",
@@ -112,13 +91,6 @@
"diff": 0.005, "diff": 0.005,
"tokens": 2474 "tokens": 2474
}, },
{
"object": "ZCL_DEMO_ABAP_NUMERIC_OP (CLAS)",
"version": "v816",
"newest": "main",
"diff": 0.003,
"tokens": 20484
},
{ {
"object": "ZDEMO_ABAP_ALV (PROG)", "object": "ZDEMO_ABAP_ALV (PROG)",
"version": "v757", "version": "v757",
@@ -127,22 +99,22 @@
"tokens": 7268 "tokens": 7268
} }
], ],
"split_docs": 43, "split_docs": 41,
"hard_line_splits": 17, "hard_line_splits": 15,
"pieces_over_limit": 0, "pieces_over_limit": 0,
"train": { "train": {
"docs": 374, "docs": 374,
"tokens": 2567180, "tokens": 2525568,
"min": 49, "min": 49,
"p50": 4904, "p50": 4770,
"p90": 15638, "p90": 15590,
"p99": 16177, "p99": 16177,
"max": 16252 "max": 16252
}, },
"valid": { "valid": {
"docs": 20, "docs": 20,
"tokens": 134816, "tokens": 135266,
"min": 92, "min": 135,
"p50": 7589, "p50": 7589,
"p90": 15317, "p90": 15317,
"p99": 16273, "p99": 16273,
@@ -150,10 +122,10 @@
}, },
"all": { "all": {
"docs": 394, "docs": 394,
"tokens": 2701996, "tokens": 2660834,
"min": 49, "min": 49,
"p50": 4904, "p50": 4797,
"p90": 15608, "p90": 15577,
"p99": 16217, "p99": 16217,
"max": 16273 "max": 16273
}, },
@@ -173,22 +145,22 @@
}, },
"share_final": { "share_final": {
"CLAS": [ "CLAS": [
1297878, 1247458,
48.0 46.9
], ],
"other": [ "other": [
101390, 110648,
3.8 4.2
], ],
"DOC": [ "DOC": [
1302728, 1302728,
48.2 49.0
] ]
}, },
"by_type": { "by_type": {
"CLAS": { "CLAS": {
"docs": 194, "docs": 187,
"tokens": 1297878 "tokens": 1247458
}, },
"mixed(4)": { "mixed(4)": {
"docs": 6, "docs": 6,
@@ -203,12 +175,12 @@
"tokens": 1792 "tokens": 1792
}, },
"BDEF": { "BDEF": {
"docs": 4, "docs": 5,
"tokens": 1069 "tokens": 1262
}, },
"DDLS": { "DDLS": {
"docs": 21, "docs": 27,
"tokens": 11549 "tokens": 20614
}, },
"mixed(3)": { "mixed(3)": {
"docs": 2, "docs": 2,
@@ -423,6 +395,17 @@
1144 1144
] ]
}, },
{
"object": [
"ZCL_DEMO_ABAP_NUMERIC_OP__V816 (CLAS)"
],
"tokens": 19769,
"pieces": 2,
"piece_tokens": [
15819,
4012
]
},
{ {
"object": [ "object": [
"05_Constructor_Expressions__P01 (DOC)" "05_Constructor_Expressions__P01 (DOC)"
@@ -434,28 +417,6 @@
2265 2265
] ]
}, },
{
"object": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V755 (CLAS)"
],
"tokens": 17237,
"pieces": 2,
"piece_tokens": [
15833,
1469
]
},
{
"object": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V756 (CLAS)"
],
"tokens": 18069,
"pieces": 2,
"piece_tokens": [
15842,
2292
]
},
{ {
"object": [ "object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)" "ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
@@ -632,6 +593,17 @@
4797 4797
] ]
}, },
{
"object": [
"ZCL_DEMO_ABAP_UNIT_TEST__V758 (CLAS)"
],
"tokens": 16394,
"pieces": 2,
"piece_tokens": [
15301,
1187
]
},
{ {
"object": [ "object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)" "ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)"
@@ -676,28 +648,6 @@
5831 5831
] ]
}, },
{
"object": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V755 (CLAS)"
],
"tokens": 17309,
"pieces": 2,
"piece_tokens": [
2124,
15283
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V756 (CLAS)"
],
"tokens": 20979,
"pieces": 2,
"piece_tokens": [
15831,
5210
]
},
{ {
"object": [ "object": [
"16_Data_Types_and_Objects__P02 (DOC)" "16_Data_Types_and_Objects__P02 (DOC)"

View File

@@ -3,8 +3,8 @@
train/.venv/bin/python train/prepare.py [--corpus PATH] [--max-len 16384] [--seed 20261003] train/.venv/bin/python train/prepare.py [--corpus PATH] [--max-len 16384] [--seed 20261003]
1. Real token counts (tokenizer of the base model). 1. Real token counts (tokenizer of the base model).
2. Version dedup: in each family keep the newest version; keep an older version only if its source differs 2. Version dedup (strict): keep main (Cloud) and the newest v* (Standard ABAP); compare each older v* with the
from the newest by more than 5 % of lines (difflib). next newer v*, keep it only if more than 5 % of lines differ (difflib).
3. Documents longer than max_len are split with the real tokenizer: classes at ENDMETHOD boundaries, markdown 3. Documents longer than max_len are split with the real tokenizer: classes at ENDMETHOD boundaries, markdown
at "##" headings (fallbacks: "###", then lines). Header lines are repeated in each piece. at "##" headings (fallbacks: "###", then lines). Header lines are repeated in each piece.
4. Split 95/5 by family (all versions and pieces of one family stay in one split). 4. Split 95/5 by family (all versions and pieces of one family stay in one split).
@@ -201,29 +201,49 @@ def main():
removed_versions, kept_versions, strict_extra = [], 0, 0 removed_versions, kept_versions, strict_extra = [], 0, 0
keep_ids = set() keep_ids = set()
n_multi = 0 n_multi = 0
other_cmp = 0
def keep(v_docs):
for d in v_docs:
keep_ids.add(d["_id"])
for f, vers in byfam.items(): for f, vers in byfam.items():
order = sorted(vers, key=lambda v: -RANK.get(v, 40)) # Strict rule (Kral 2026-10-03): main (ABAP Cloud) and the newest v* (Standard ABAP) are always kept.
newest = order[0] # Each older v* is compared with the next newer v* (not with main) and kept only if > diff of lines differ.
for d in vers[newest]: # Other branches (oo_patterns, rap, unit_tests) are compared with main as before.
keep_ids.add(d["_id"]) vs = sorted((v for v in vers if re.fullmatch(r"v\d+", v)), key=lambda v: -int(v[1:]))
kept_versions += 1 others = [v for v in vers if v != "main" and v not in vs]
if len(order) > 1: if len(vers) > 1:
n_multi += 1 n_multi += 1
base = [x for d in vers[newest] for x in norm_lines(d["text"])] lines_of = lambda v: [x for d in vers[v] for x in norm_lines(d["text"])]
kept_lines = [base] if "main" in vers:
for v in order[1:]: keep(vers["main"])
lines = [x for d in vers[v] for x in norm_lines(d["text"])] kept_versions += 1
frac = diff_fraction(base, lines) if vs:
if frac > a.diff: keep(vers[vs[0]])
kept_versions += 1
for i in range(1, len(vs)):
newer = vs[i - 1]
frac = diff_fraction(lines_of(newer), lines_of(vs[i]))
if frac > a.diff:
keep(vers[vs[i]])
kept_versions += 1 kept_versions += 1
for d in vers[v]:
keep_ids.add(d["_id"])
# info only: would a stricter rule (compare with all kept versions) remove it?
if min(diff_fraction(k, lines) for k in kept_lines) <= a.diff:
strict_extra += 1
kept_lines.append(lines)
else: else:
removed_versions.append({"object": f, "version": v, "newest": newest, "diff": round(frac, 3), removed_versions.append({"object": f, "version": vs[i], "newest": newer, "diff": round(frac, 3),
"tokens": sum(d["real_tokens"] for d in vers[vs[i]])})
for v in others:
ref = "main" if "main" in vers else (vs[0] if vs else None)
if ref is None:
keep(vers[v])
kept_versions += 1
continue
other_cmp += 1
frac = diff_fraction(lines_of(ref), lines_of(v))
if frac > a.diff:
keep(vers[v])
kept_versions += 1
else:
removed_versions.append({"object": f, "version": v, "newest": ref, "diff": round(frac, 3),
"tokens": sum(d["real_tokens"] for d in vers[v])}) "tokens": sum(d["real_tokens"] for d in vers[v])})
after_dedup = [d for d in docs if d["_id"] in keep_ids] after_dedup = [d for d in docs if d["_id"] in keep_ids]
real1 = sum(d["real_tokens"] for d in after_dedup) real1 = sum(d["real_tokens"] for d in after_dedup)
@@ -282,7 +302,7 @@ def main():
"records": len(docs), "tokens_before_dedup": real0, "tokens_after_dedup": real1, "records": len(docs), "tokens_before_dedup": real0, "tokens_after_dedup": real1,
"records_after_dedup": len(after_dedup), "tokens_after_split": real2, "records_after_dedup": len(after_dedup), "tokens_after_split": real2,
"objects_with_versions": n_multi, "families": len(byfam), "versions_kept": kept_versions, "objects_with_versions": n_multi, "families": len(byfam), "versions_kept": kept_versions,
"versions_removed": len(removed_versions), "strict_rule_would_remove_more": strict_extra, "versions_removed": len(removed_versions), "other_branch_comparisons": other_cmp,
"removed_versions": removed_versions, "split_docs": len(split_info), "hard_line_splits": hard_total, "removed_versions": removed_versions, "split_docs": len(split_info), "hard_line_splits": hard_total,
"pieces_over_limit": len(over), "train": summ(train), "valid": summ(valid), "all": summ(kept), "pieces_over_limit": len(over), "train": summ(train), "valid": summ(valid), "all": summ(kept),
"share_before_dedup": share([{"types": d["types"], "tokens": d["real_tokens"]} for d in docs]), "share_before_dedup": share([{"types": d["types"], "tokens": d["real_tokens"]} for d in docs]),
@@ -293,13 +313,13 @@ def main():
L = ["# Stage 1 data report", "", L = ["# Stage 1 data report", "",
f"Corpus: `{a.corpus}` (SAP-samples/abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md).", f"Corpus: `{a.corpus}` (SAP-samples/abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md).",
f"Tokenizer: `{a.model}`. Seed {a.seed}. max_seq_length {a.max_len}.", "", f"Tokenizer: `{a.model}`. Seed {a.seed}. max_seq_length {a.max_len}.", "",
"## Version dedup (keep the newest; keep an older version only if it differs by more than " "## Version dedup (strict rule; an older version is kept only if it differs by more than "
f"{int(a.diff * 100)} % of lines)", "", f"{int(a.diff * 100)} % of lines)", "",
f"- Object families: {len(byfam)}; with more than one version: {n_multi}.", f"- Object families: {len(byfam)}; with more than one version: {n_multi}.",
f"- Versions kept: {kept_versions}. Versions removed: {len(removed_versions)}.", f"- Versions kept: {kept_versions}. Versions removed: {len(removed_versions)}.",
f"- Records: {len(docs)} before, {len(after_dedup)} after.", f"- Records: {len(docs)} before, {len(after_dedup)} after.",
f"- Tokens before: {real0} (chars/4 estimate {est}). After dedup: {real1}. After splitting: {real2}.", f"- Tokens before: {real0} (chars/4 estimate {est}). After dedup: {real1}. After splitting: {real2}.",
f"- Info: a stricter rule (compare also with the other kept versions) would remove {strict_extra} more kept versions.", f"- Rule: main (Cloud) and the newest v* (Standard ABAP) always kept; each older v* compared with the next newer v*; other branches ({other_cmp} comparisons) compared with main.",
"", "## Splitting of documents over the limit", "", "", "## Splitting of documents over the limit", "",
f"- Documents split: {len(split_info)}; pieces: {sum(s['pieces'] for s in split_info)}; hard line splits " f"- Documents split: {len(split_info)}; pieces: {sum(s['pieces'] for s in split_info)}; hard line splits "
f"(a block was too big): {hard_total}; pieces still over the limit: {len(over)}.", "", f"(a block was too big): {hard_total}; pieces still over the limit: {len(over)}.", "",