Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 22:09:15 +02:00
parent e083c9ca13
commit 0677d035da
5 changed files with 1014 additions and 421 deletions

View File

@@ -13,8 +13,9 @@ Stage 2 (SFT on teacher trajectories) is not part of this task.
Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03): Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03):
SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`). SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`).
Real numbers from Step 1 (`train/data/report.md`): 370 documents, 2,806,501 tokens (Qwen tokenizer; Real numbers from Step 1 (`train/data/report.md`): 370 records, 2,806,501 tokens (Qwen tokenizer;
the chars/4 estimate is 2.64M). The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line: the chars/4 estimate is 2.64M). After version dedup and splitting: train 374 documents / 2.57M tokens,
valid 20 documents / 135k tokens, 748 iterations for 2 epochs. The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line:
`{"text", "objects", "types", "language_version", "package_path", `{"text", "objects", "types", "language_version", "package_path",
"grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`. "grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`.

View File

@@ -1,6 +1,6 @@
# Stage 1 state # Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 19:10. Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:30.
## Done ## Done
@@ -15,10 +15,14 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`, - Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`). run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:00): `train/prepare.py`, report `train/data/report.md`. Corpus replaced by - Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): `train/prepare.py`, report `train/data/report.md`.
SAP-samples/abap-cheat-sheets (Apache-2.0): 370 docs, 2.81M real tokens. 45 docs (840k tokens, 30 %) are longer Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
than 16384 and removed. Train 309 docs / 1.89M tokens, valid 16 docs / 80k tokens (split by family, seed Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept;
20261003), 618 iterations for 2 epochs. `test.jsonl` = copy of valid (needed by `mlx_lm.lora --test`). tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces
(classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece
over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %.
Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**618 -> 748 iterations for 2 epochs.** `test.jsonl` = copy of valid.
## Running (detached) ## Running (detached)
@@ -32,7 +36,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json` 1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit. (`t01_test_budget40` holds the T01 test result). Commit.
2. Decisions for Kral on the data (see report): the 45 removed docs, DOC share, valid size. 2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the 3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`. options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20

View File

@@ -1,98 +1,125 @@
# Stage 1 data report # Stage 1 data report
Corpus: `/Users/erhankeseli/projects/abap-llm/corpus/corpus.jsonl` (source abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md). Corpus: `/Users/erhankeseli/projects/abap-llm/corpus/corpus.jsonl` (SAP-samples/abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md).
Tokenizer: `/Users/erhankeseli/models/Qwen3.8-27B-4bit`. Seed 20261003. max_seq_length 16384. Tokenizer: `/Users/erhankeseli/models/Qwen3.8-27B-4bit`. Seed 20261003. max_seq_length 16384.
Split by family (release versions and parts of one document stay together): 180 families.
## Version dedup (keep the newest; keep an older version only if it differs by more than 5 % of lines)
- Object families: 269; with more than one version: 39.
- Versions kept: 319. Versions removed: 16.
- Records: 370 before, 351 after.
- Tokens before: 2806501 (chars/4 estimate 2636112). After dedup: 2699110. After splitting: 2701996.
- Info: a stricter rule (compare also with the other kept versions) would remove 8 more kept versions.
## Splitting of documents over the limit
- Documents split: 43; pieces: 86; hard line splits (a block was too big): 17; pieces still over the limit: 0.
## Result
| | docs | tokens | | | docs | tokens |
|---|---|---| |---|---|---|
| corpus | 370 | 2806501 (estimate chars/4: 2636112, ratio 1.065) | | train | 374 | 2567180 |
| removed (> 16384) | 45 | 840136 | | valid (= test file) | 20 | 134816 |
| train | 309 | 1886475 |
| valid (= test file) | 16 | 79890 |
Iterations for 2 epochs (batch 1): 618. Split by family: 184 families. **Iterations for 2 epochs (batch 1): 748.**
## Token distribution per document (real tokenizer) ## Token share by type
| | before dedup | final (kept pieces) |
|---|---|---|
| DOC | 1301698 (46.4 %) | 1302728 (48.2 %) |
| CLAS | 1378559 (49.1 %) | 1297878 (48.0 %) |
| other | 126244 (4.5 %) | 101390 (3.8 %) |
## Token distribution per document (final)
| set | min | p50 | p90 | p99 | max | | set | min | p50 | p90 | p99 | max |
|---|---|---|---|---|---| |---|---|---|---|---|---|
| all | 49 | 5195 | 16923 | 21483 | 31552 | | all | 49 | 4904 | 15608 | 16217 | 16273 |
| kept | 49 | 4027 | 14998 | 16217 | 16273 | | train | 49 | 4904 | 15638 | 16177 | 16252 |
| train | 49 | 4403 | 14998 | 16217 | 16273 | | valid | 92 | 7589 | 15317 | 16273 | 16273 |
| valid | 130 | 3392 | 16027 | 16129 | 16129 |
| tokens per document up to | docs | ## By object type (final)
|---|---|
| 512 | 48 |
| 1024 | 25 |
| 2048 | 40 |
| 4096 | 51 |
| 8192 | 53 |
| 16384 | 108 |
| more | 45 |
## By object type (kept documents)
| types | docs | tokens | | types | docs | tokens |
|---|---|---| |---|---|---|
| DOC | 106 | 1012152 | | DOC | 138 | 1302728 |
| CLAS | 147 | 827969 | | CLAS | 194 | 1297878 |
| PROG | 13 | 55365 | | PROG | 11 | 40071 |
| mixed(4) | 6 | 28405 | | mixed(4) | 6 | 28405 |
| DDLS | 27 | 20614 | | DDLS | 21 | 11549 |
| mixed(3) | 2 | 6667 | | mixed(3) | 2 | 6667 |
| DDLX | 6 | 5003 | | DDLX | 6 | 5003 |
| CLAS+DDLS | 1 | 4563 | | CLAS+DDLS | 1 | 4563 |
| FUGR | 2 | 2026 | | FUGR | 2 | 2026 |
| INTF | 7 | 1792 | | INTF | 7 | 1792 |
| BDEF | 6 | 1564 | | BDEF | 4 | 1069 |
| DCLS | 2 | 245 | | DCLS | 2 | 245 |
## Removed documents ## Removed versions
- ZCL_DEMO_ABAP_OBJECTS (CLAS): 17042 tokens | object | version | newest | diff | tokens |
- ZCL_DEMO_ABAP_STRUCTURES (CLAS): 16628 tokens |---|---|---|---|---|
- 32_Performance_Notes__P03 (DOC): 16742 tokens | ZBP_DEMO_ABAP_RAP_RO_U (CLAS) | v758 | main | 0.015 | 10696 |
- ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS): 19117 tokens | ZDEMO_ABAP_RAP_RO_M (BDEF) | v756 | main | 0.037 | 259 |
- ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS): 18883 tokens | ZDEMO_ABAP_RAP_RO_U (BDEF) | v756 | main | 0.043 | 236 |
- ZCL_DEMO_ABAP_DYNAMIC_PROG__V755 (CLAS): 19801 tokens | ZDEMO_ABAP_CDS_VE_ASSOC_E (DDLS) | v757 | main | 0.026 | 542 |
- 23_Date_and_Time (DOC): 20347 tokens | ZDEMO_ABAP_CDS_VE_JOINS (DDLS) | v757 | main | 0.009 | 1827 |
- ZCL_DEMO_ABAP_INTERNAL_TABLES__V757 (CLAS): 17477 tokens | ZDEMO_ABAP_CDS_VE_AGG_EXP (DDLS) | v757 | main | 0.019 | 804 |
- 07_String_Processing__P02 (DOC): 17916 tokens | ZDEMO_ABAP_CDS_VE_SEL (DDLS) | v757 | main | 0.029 | 3273 |
- ZCL_DEMO_ABAP_CONSTRUCTOR_EXPR (CLAS): 17368 tokens | ZDEMO_ABAP_CDS_VE_ASSOC (DDLS) | v757 | main | 0.008 | 2451 |
- ZCL_DEMO_ABAP_DYNAMIC_PROG__V758 (CLAS): 31552 tokens | ZDEMO_ABAP_FLI_VE (DDLS) | v816 | main | 0.0 | 168 |
- ZCL_DEMO_ABAP_STRING_PROC__V757 (CLAS): 19499 tokens | ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS) | v816 | main | 0.009 | 27407 |
- 04_ABAP_Object_Orientation__P06 (DOC): 16694 tokens | ZCL_DEMO_ABAP_DISPLAY (CLAS) | v755 | v757 | 0.02 | 1677 |
- ZCL_DEMO_ABAP_DTYPE_DOBJ__V758 (CLAS): 18368 tokens | ZCL_DEMO_ABAP_UNIT_TEST (CLAS) | v758 | main | 0.034 | 19799 |
- 06_Dynamic_Programming__P07 (DOC): 23994 tokens | ZDEMO_ABAP_DYNPRO (PROG) | v757 | v816 | 0.001 | 8026 |
- 28_Regular_Expressions (DOC): 17900 tokens | ZBP_DEMO_ABAP_RAP_DRAFT_M (CLAS) | v757 | v758 | 0.005 | 2474 |
- 06_Dynamic_Programming__P06 (DOC): 16923 tokens | ZCL_DEMO_ABAP_NUMERIC_OP (CLAS) | v816 | main | 0.003 | 20484 |
- ZCL_DEMO_ABAP_NUMERIC_OP__V816 (CLAS): 19769 tokens | ZDEMO_ABAP_ALV (PROG) | v757 | v816 | 0.002 | 7268 |
- 05_Constructor_Expressions__P01 (DOC): 17349 tokens
- ZCL_DEMO_ABAP_DTYPE_DOBJ__V755 (CLAS): 17237 tokens ## Split documents
- ZCL_DEMO_ABAP_DTYPE_DOBJ__V756 (CLAS): 18069 tokens
- ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS): 17923 tokens - ['ZCL_DEMO_ABAP_OBJECTS (CLAS)']: 17042 tokens -> 2 pieces [14777, 2342]
- ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS): 16924 tokens - ['ZCL_DEMO_ABAP_STRUCTURES (CLAS)']: 16628 tokens -> 2 pieces [14200, 2503]
- ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS): 16770 tokens - ['32_Performance_Notes__P03 (DOC)']: 16742 tokens -> 2 pieces [15846, 961]
- 14_ABAP_Unit_Tests__P03 (DOC): 18549 tokens - ['ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)']: 19117 tokens -> 2 pieces [14425, 4770]
- ZCL_DEMO_ABAP_STRING_PROC__V758 (CLAS): 22096 tokens - ['ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)']: 18883 tokens -> 2 pieces [14607, 4324]
- 22_Released_ABAP_Classes__P01 (DOC): 19022 tokens - ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V755 (CLAS)']: 19801 tokens -> 2 pieces [15826, 4036]
- 21_XML_JSON__P01 (DOC): 16625 tokens - ['23_Date_and_Time (DOC)']: 20347 tokens -> 2 pieces [10310, 10098]
- 29_Numeric_Operations__P01 (DOC): 17805 tokens - ['ZCL_DEMO_ABAP_INTERNAL_TABLES__V757 (CLAS)']: 17477 tokens -> 2 pieces [2125, 15451]
- 03_ABAP_SQL__P01 (DOC): 17305 tokens - ['07_String_Processing__P02 (DOC)']: 17916 tokens -> 2 pieces [11096, 6879]
- ZCL_DEMO_ABAP_DTYPE_DOBJ__V757 (CLAS): 18308 tokens - ['ZCL_DEMO_ABAP_CONSTRUCTOR_EXPR (CLAS)']: 17368 tokens -> 2 pieces [14031, 3415]
- 14_ABAP_Unit_Tests__P04 (DOC): 17633 tokens - ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V758 (CLAS)']: 31552 tokens -> 2 pieces [15833, 15780]
- 24_Builtin_Functions (DOC): 16851 tokens - ['ZCL_DEMO_ABAP_STRING_PROC__V757 (CLAS)']: 19499 tokens -> 2 pieces [15831, 3729]
- ZCL_DEMO_ABAP_NUMERIC_OP (CLAS): 19740 tokens - ['04_ABAP_Object_Orientation__P06 (DOC)']: 16694 tokens -> 2 pieces [9045, 7716]
- ZCL_DEMO_ABAP_INTERNAL_TABLES__V758 (CLAS): 17355 tokens - ['ZCL_DEMO_ABAP_DTYPE_DOBJ__V758 (CLAS)']: 18368 tokens -> 2 pieces [15833, 2600]
- ZCL_DEMO_ABAP_BUILTIN_FUNC (CLAS): 20446 tokens - ['06_Dynamic_Programming__P07 (DOC)']: 23994 tokens -> 2 pieces [15836, 8221]
- ZCL_DEMO_ABAP_STRING_PROC (CLAS): 20311 tokens - ['28_Regular_Expressions (DOC)']: 17900 tokens -> 2 pieces [15031, 2930]
- ZCL_DEMO_ABAP_UNIT_TEST__V758 (CLAS): 16394 tokens - ['06_Dynamic_Programming__P06 (DOC)']: 16923 tokens -> 2 pieces [15842, 1144]
- ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS): 21483 tokens - ['05_Constructor_Expressions__P01 (DOC)']: 17349 tokens -> 2 pieces [15147, 2265]
- ZCL_DEMO_ABAP_DATE_TIME (CLAS): 18634 tokens - ['ZCL_DEMO_ABAP_DTYPE_DOBJ__V755 (CLAS)']: 17237 tokens -> 2 pieces [15833, 1469]
- ZCL_DEMO_ABAP_XML_JSON (CLAS): 16825 tokens - ['ZCL_DEMO_ABAP_DTYPE_DOBJ__V756 (CLAS)']: 18069 tokens -> 2 pieces [15842, 2292]
- ZCL_DEMO_ABAP_SQL (CLAS): 18283 tokens - ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 17923 tokens -> 2 pieces [15329, 2672]
- ZCL_DEMO_ABAP_INTERNAL_TABLES__V755 (CLAS): 17309 tokens - ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 16924 tokens -> 2 pieces [15749, 1223]
- ZCL_DEMO_ABAP_DYNAMIC_PROG__V756 (CLAS): 20979 tokens - ['ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)']: 16770 tokens -> 2 pieces [14210, 2608]
- 16_Data_Types_and_Objects__P02 (DOC): 17891 tokens - ['14_ABAP_Unit_Tests__P03 (DOC)']: 18549 tokens -> 2 pieces [15834, 2782]
- ['ZCL_DEMO_ABAP_STRING_PROC__V758 (CLAS)']: 22096 tokens -> 2 pieces [15814, 6342]
- ['22_Released_ABAP_Classes__P01 (DOC)']: 19022 tokens -> 2 pieces [14661, 4430]
- ['21_XML_JSON__P01 (DOC)']: 16625 tokens -> 2 pieces [9855, 6839]
- ['29_Numeric_Operations__P01 (DOC)']: 17805 tokens -> 2 pieces [15849, 2025]
- ['03_ABAP_SQL__P01 (DOC)']: 17305 tokens -> 2 pieces [14168, 3198]
- ['ZCL_DEMO_ABAP_DTYPE_DOBJ__V757 (CLAS)']: 18308 tokens -> 2 pieces [15842, 2531]
- ['14_ABAP_Unit_Tests__P04 (DOC)']: 17633 tokens -> 2 pieces [15834, 1866]
- ['24_Builtin_Functions (DOC)']: 16851 tokens -> 2 pieces [12541, 4365]
- ['ZCL_DEMO_ABAP_NUMERIC_OP (CLAS)']: 19740 tokens -> 2 pieces [15825, 3964]
- ['ZCL_DEMO_ABAP_INTERNAL_TABLES__V758 (CLAS)']: 17355 tokens -> 2 pieces [15832, 1584]
- ['ZCL_DEMO_ABAP_BUILTIN_FUNC (CLAS)']: 20446 tokens -> 2 pieces [15454, 5071]
- ['ZCL_DEMO_ABAP_STRING_PROC (CLAS)']: 20311 tokens -> 2 pieces [15590, 4797]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)']: 21483 tokens -> 2 pieces [15815, 5730]
- ['ZCL_DEMO_ABAP_DATE_TIME (CLAS)']: 18634 tokens -> 2 pieces [15577, 3138]
- ['ZCL_DEMO_ABAP_XML_JSON (CLAS)']: 16825 tokens -> 2 pieces [15753, 1154]
- ['ZCL_DEMO_ABAP_SQL (CLAS)']: 18283 tokens -> 2 pieces [12526, 5831]
- ['ZCL_DEMO_ABAP_INTERNAL_TABLES__V755 (CLAS)']: 17309 tokens -> 2 pieces [2124, 15283]
- ['ZCL_DEMO_ABAP_DYNAMIC_PROG__V756 (CLAS)']: 20979 tokens -> 2 pieces [15831, 5210]
- ['16_Data_Types_and_Objects__P02 (DOC)']: 17891 tokens -> 2 pieces [14421, 3541]

View File

@@ -3,348 +3,220 @@
"seed": 20261003, "seed": 20261003,
"max_len": 16384, "max_len": 16384,
"estimate_tokens": 2636112, "estimate_tokens": 2636112,
"real_tokens": 2806501, "records": 370,
"ratio_real_to_estimate": 1.065, "tokens_before_dedup": 2806501,
"all": { "tokens_after_dedup": 2699110,
"docs": 370, "records_after_dedup": 351,
"tokens": 2806501, "tokens_after_split": 2701996,
"min": 49, "objects_with_versions": 39,
"p50": 5195, "families": 269,
"p90": 16923, "versions_kept": 319,
"p99": 21483, "versions_removed": 16,
"max": 31552 "strict_rule_would_remove_more": 8,
}, "removed_versions": [
"kept": {
"docs": 325,
"tokens": 1966365,
"min": 49,
"p50": 4027,
"p90": 14998,
"p99": 16217,
"max": 16273
},
"train": {
"docs": 309,
"tokens": 1886475,
"min": 49,
"p50": 4403,
"p90": 14998,
"p99": 16217,
"max": 16273
},
"valid": {
"docs": 16,
"tokens": 79890,
"min": 130,
"p50": 3392,
"p90": 16027,
"p99": 16129,
"max": 16129
},
"removed": [
{ {
"objects": [ "object": "ZBP_DEMO_ABAP_RAP_RO_U (CLAS)",
"ZCL_DEMO_ABAP_OBJECTS (CLAS)" "version": "v758",
], "newest": "main",
"tokens": 17042 "diff": 0.015,
"tokens": 10696
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_RAP_RO_M (BDEF)",
"ZCL_DEMO_ABAP_STRUCTURES (CLAS)" "version": "v756",
], "newest": "main",
"tokens": 16628 "diff": 0.037,
"tokens": 259
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_RAP_RO_U (BDEF)",
"32_Performance_Notes__P03 (DOC)" "version": "v756",
], "newest": "main",
"tokens": 16742 "diff": 0.043,
"tokens": 236
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_CDS_VE_ASSOC_E (DDLS)",
"ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)" "version": "v757",
], "newest": "main",
"tokens": 19117 "diff": 0.026,
"tokens": 542
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_CDS_VE_JOINS (DDLS)",
"ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)" "version": "v757",
], "newest": "main",
"tokens": 18883 "diff": 0.009,
"tokens": 1827
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_CDS_VE_AGG_EXP (DDLS)",
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V755 (CLAS)" "version": "v757",
], "newest": "main",
"tokens": 19801 "diff": 0.019,
"tokens": 804
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_CDS_VE_SEL (DDLS)",
"23_Date_and_Time (DOC)" "version": "v757",
], "newest": "main",
"tokens": 20347 "diff": 0.029,
"tokens": 3273
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_CDS_VE_ASSOC (DDLS)",
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V757 (CLAS)" "version": "v757",
], "newest": "main",
"tokens": 17477 "diff": 0.008,
"tokens": 2451
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_FLI_VE (DDLS)",
"07_String_Processing__P02 (DOC)" "version": "v816",
], "newest": "main",
"tokens": 17916 "diff": 0.0,
"tokens": 168
}, },
{ {
"objects": [ "object": "ZCL_DEMO_ABAP_DTYPE_DOBJ (CLAS)",
"ZCL_DEMO_ABAP_CONSTRUCTOR_EXPR (CLAS)" "version": "v816",
], "newest": "main",
"tokens": 17368 "diff": 0.009,
"tokens": 27407
}, },
{ {
"objects": [ "object": "ZCL_DEMO_ABAP_DISPLAY (CLAS)",
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V758 (CLAS)" "version": "v755",
], "newest": "v757",
"tokens": 31552 "diff": 0.02,
"tokens": 1677
}, },
{ {
"objects": [ "object": "ZCL_DEMO_ABAP_UNIT_TEST (CLAS)",
"ZCL_DEMO_ABAP_STRING_PROC__V757 (CLAS)" "version": "v758",
], "newest": "main",
"tokens": 19499 "diff": 0.034,
"tokens": 19799
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_DYNPRO (PROG)",
"04_ABAP_Object_Orientation__P06 (DOC)" "version": "v757",
], "newest": "v816",
"tokens": 16694 "diff": 0.001,
"tokens": 8026
}, },
{ {
"objects": [ "object": "ZBP_DEMO_ABAP_RAP_DRAFT_M (CLAS)",
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V758 (CLAS)" "version": "v757",
], "newest": "v758",
"tokens": 18368 "diff": 0.005,
"tokens": 2474
}, },
{ {
"objects": [ "object": "ZCL_DEMO_ABAP_NUMERIC_OP (CLAS)",
"06_Dynamic_Programming__P07 (DOC)" "version": "v816",
], "newest": "main",
"tokens": 23994 "diff": 0.003,
"tokens": 20484
}, },
{ {
"objects": [ "object": "ZDEMO_ABAP_ALV (PROG)",
"28_Regular_Expressions (DOC)" "version": "v757",
], "newest": "v816",
"tokens": 17900 "diff": 0.002,
}, "tokens": 7268
{
"objects": [
"06_Dynamic_Programming__P06 (DOC)"
],
"tokens": 16923
},
{
"objects": [
"ZCL_DEMO_ABAP_NUMERIC_OP__V816 (CLAS)"
],
"tokens": 19769
},
{
"objects": [
"05_Constructor_Expressions__P01 (DOC)"
],
"tokens": 17349
},
{
"objects": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V755 (CLAS)"
],
"tokens": 17237
},
{
"objects": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V756 (CLAS)"
],
"tokens": 18069
},
{
"objects": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
],
"tokens": 17923
},
{
"objects": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
],
"tokens": 16924
},
{
"objects": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
],
"tokens": 16770
},
{
"objects": [
"14_ABAP_Unit_Tests__P03 (DOC)"
],
"tokens": 18549
},
{
"objects": [
"ZCL_DEMO_ABAP_STRING_PROC__V758 (CLAS)"
],
"tokens": 22096
},
{
"objects": [
"22_Released_ABAP_Classes__P01 (DOC)"
],
"tokens": 19022
},
{
"objects": [
"21_XML_JSON__P01 (DOC)"
],
"tokens": 16625
},
{
"objects": [
"29_Numeric_Operations__P01 (DOC)"
],
"tokens": 17805
},
{
"objects": [
"03_ABAP_SQL__P01 (DOC)"
],
"tokens": 17305
},
{
"objects": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V757 (CLAS)"
],
"tokens": 18308
},
{
"objects": [
"14_ABAP_Unit_Tests__P04 (DOC)"
],
"tokens": 17633
},
{
"objects": [
"24_Builtin_Functions (DOC)"
],
"tokens": 16851
},
{
"objects": [
"ZCL_DEMO_ABAP_NUMERIC_OP (CLAS)"
],
"tokens": 19740
},
{
"objects": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V758 (CLAS)"
],
"tokens": 17355
},
{
"objects": [
"ZCL_DEMO_ABAP_BUILTIN_FUNC (CLAS)"
],
"tokens": 20446
},
{
"objects": [
"ZCL_DEMO_ABAP_STRING_PROC (CLAS)"
],
"tokens": 20311
},
{
"objects": [
"ZCL_DEMO_ABAP_UNIT_TEST__V758 (CLAS)"
],
"tokens": 16394
},
{
"objects": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)"
],
"tokens": 21483
},
{
"objects": [
"ZCL_DEMO_ABAP_DATE_TIME (CLAS)"
],
"tokens": 18634
},
{
"objects": [
"ZCL_DEMO_ABAP_XML_JSON (CLAS)"
],
"tokens": 16825
},
{
"objects": [
"ZCL_DEMO_ABAP_SQL (CLAS)"
],
"tokens": 18283
},
{
"objects": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V755 (CLAS)"
],
"tokens": 17309
},
{
"objects": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V756 (CLAS)"
],
"tokens": 20979
},
{
"objects": [
"16_Data_Types_and_Objects__P02 (DOC)"
],
"tokens": 17891
} }
], ],
"split_docs": 43,
"hard_line_splits": 17,
"pieces_over_limit": 0,
"train": {
"docs": 374,
"tokens": 2567180,
"min": 49,
"p50": 4904,
"p90": 15638,
"p99": 16177,
"max": 16252
},
"valid": {
"docs": 20,
"tokens": 134816,
"min": 92,
"p50": 7589,
"p90": 15317,
"p99": 16273,
"max": 16273
},
"all": {
"docs": 394,
"tokens": 2701996,
"min": 49,
"p50": 4904,
"p90": 15608,
"p99": 16217,
"max": 16273
},
"share_before_dedup": {
"CLAS": [
1378559,
49.1
],
"other": [
126244,
4.5
],
"DOC": [
1301698,
46.4
]
},
"share_final": {
"CLAS": [
1297878,
48.0
],
"other": [
101390,
3.8
],
"DOC": [
1302728,
48.2
]
},
"by_type": { "by_type": {
"CLAS": { "CLAS": {
"docs": 147, "docs": 194,
"tokens": 827969 "tokens": 1297878
}, },
"mixed(4)": { "mixed(4)": {
"docs": 6, "docs": 6,
"tokens": 28405 "tokens": 28405
}, },
"DOC": { "DOC": {
"docs": 106, "docs": 138,
"tokens": 1012152 "tokens": 1302728
}, },
"INTF": { "INTF": {
"docs": 7, "docs": 7,
"tokens": 1792 "tokens": 1792
}, },
"BDEF": { "BDEF": {
"docs": 6, "docs": 4,
"tokens": 1564 "tokens": 1069
}, },
"DDLS": { "DDLS": {
"docs": 27, "docs": 21,
"tokens": 20614 "tokens": 11549
}, },
"mixed(3)": { "mixed(3)": {
"docs": 2, "docs": 2,
"tokens": 6667 "tokens": 6667
}, },
"PROG": { "PROG": {
"docs": 13, "docs": 11,
"tokens": 55365 "tokens": 40071
}, },
"DDLX": { "DDLX": {
"docs": 6, "docs": 6,
@@ -363,8 +235,480 @@
"tokens": 2026 "tokens": 2026
} }
}, },
"source": [ "split_info": [
"abap-cheat-sheets" {
"object": [
"ZCL_DEMO_ABAP_OBJECTS (CLAS)"
], ],
"iterations_2_epochs": 618 "tokens": 17042,
"pieces": 2,
"piece_tokens": [
14777,
2342
]
},
{
"object": [
"ZCL_DEMO_ABAP_STRUCTURES (CLAS)"
],
"tokens": 16628,
"pieces": 2,
"piece_tokens": [
14200,
2503
]
},
{
"object": [
"32_Performance_Notes__P03 (DOC)"
],
"tokens": 16742,
"pieces": 2,
"piece_tokens": [
15846,
961
]
},
{
"object": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)"
],
"tokens": 19117,
"pieces": 2,
"piece_tokens": [
14425,
4770
]
},
{
"object": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES (CLAS)"
],
"tokens": 18883,
"pieces": 2,
"piece_tokens": [
14607,
4324
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V755 (CLAS)"
],
"tokens": 19801,
"pieces": 2,
"piece_tokens": [
15826,
4036
]
},
{
"object": [
"23_Date_and_Time (DOC)"
],
"tokens": 20347,
"pieces": 2,
"piece_tokens": [
10310,
10098
]
},
{
"object": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V757 (CLAS)"
],
"tokens": 17477,
"pieces": 2,
"piece_tokens": [
2125,
15451
]
},
{
"object": [
"07_String_Processing__P02 (DOC)"
],
"tokens": 17916,
"pieces": 2,
"piece_tokens": [
11096,
6879
]
},
{
"object": [
"ZCL_DEMO_ABAP_CONSTRUCTOR_EXPR (CLAS)"
],
"tokens": 17368,
"pieces": 2,
"piece_tokens": [
14031,
3415
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V758 (CLAS)"
],
"tokens": 31552,
"pieces": 2,
"piece_tokens": [
15833,
15780
]
},
{
"object": [
"ZCL_DEMO_ABAP_STRING_PROC__V757 (CLAS)"
],
"tokens": 19499,
"pieces": 2,
"piece_tokens": [
15831,
3729
]
},
{
"object": [
"04_ABAP_Object_Orientation__P06 (DOC)"
],
"tokens": 16694,
"pieces": 2,
"piece_tokens": [
9045,
7716
]
},
{
"object": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V758 (CLAS)"
],
"tokens": 18368,
"pieces": 2,
"piece_tokens": [
15833,
2600
]
},
{
"object": [
"06_Dynamic_Programming__P07 (DOC)"
],
"tokens": 23994,
"pieces": 2,
"piece_tokens": [
15836,
8221
]
},
{
"object": [
"28_Regular_Expressions (DOC)"
],
"tokens": 17900,
"pieces": 2,
"piece_tokens": [
15031,
2930
]
},
{
"object": [
"06_Dynamic_Programming__P06 (DOC)"
],
"tokens": 16923,
"pieces": 2,
"piece_tokens": [
15842,
1144
]
},
{
"object": [
"05_Constructor_Expressions__P01 (DOC)"
],
"tokens": 17349,
"pieces": 2,
"piece_tokens": [
15147,
2265
]
},
{
"object": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V755 (CLAS)"
],
"tokens": 17237,
"pieces": 2,
"piece_tokens": [
15833,
1469
]
},
{
"object": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V756 (CLAS)"
],
"tokens": 18069,
"pieces": 2,
"piece_tokens": [
15842,
2292
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
],
"tokens": 17923,
"pieces": 2,
"piece_tokens": [
15329,
2672
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
],
"tokens": 16924,
"pieces": 2,
"piece_tokens": [
15749,
1223
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG (CLAS)"
],
"tokens": 16770,
"pieces": 2,
"piece_tokens": [
14210,
2608
]
},
{
"object": [
"14_ABAP_Unit_Tests__P03 (DOC)"
],
"tokens": 18549,
"pieces": 2,
"piece_tokens": [
15834,
2782
]
},
{
"object": [
"ZCL_DEMO_ABAP_STRING_PROC__V758 (CLAS)"
],
"tokens": 22096,
"pieces": 2,
"piece_tokens": [
15814,
6342
]
},
{
"object": [
"22_Released_ABAP_Classes__P01 (DOC)"
],
"tokens": 19022,
"pieces": 2,
"piece_tokens": [
14661,
4430
]
},
{
"object": [
"21_XML_JSON__P01 (DOC)"
],
"tokens": 16625,
"pieces": 2,
"piece_tokens": [
9855,
6839
]
},
{
"object": [
"29_Numeric_Operations__P01 (DOC)"
],
"tokens": 17805,
"pieces": 2,
"piece_tokens": [
15849,
2025
]
},
{
"object": [
"03_ABAP_SQL__P01 (DOC)"
],
"tokens": 17305,
"pieces": 2,
"piece_tokens": [
14168,
3198
]
},
{
"object": [
"ZCL_DEMO_ABAP_DTYPE_DOBJ__V757 (CLAS)"
],
"tokens": 18308,
"pieces": 2,
"piece_tokens": [
15842,
2531
]
},
{
"object": [
"14_ABAP_Unit_Tests__P04 (DOC)"
],
"tokens": 17633,
"pieces": 2,
"piece_tokens": [
15834,
1866
]
},
{
"object": [
"24_Builtin_Functions (DOC)"
],
"tokens": 16851,
"pieces": 2,
"piece_tokens": [
12541,
4365
]
},
{
"object": [
"ZCL_DEMO_ABAP_NUMERIC_OP (CLAS)"
],
"tokens": 19740,
"pieces": 2,
"piece_tokens": [
15825,
3964
]
},
{
"object": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V758 (CLAS)"
],
"tokens": 17355,
"pieces": 2,
"piece_tokens": [
15832,
1584
]
},
{
"object": [
"ZCL_DEMO_ABAP_BUILTIN_FUNC (CLAS)"
],
"tokens": 20446,
"pieces": 2,
"piece_tokens": [
15454,
5071
]
},
{
"object": [
"ZCL_DEMO_ABAP_STRING_PROC (CLAS)"
],
"tokens": 20311,
"pieces": 2,
"piece_tokens": [
15590,
4797
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V757 (CLAS)"
],
"tokens": 21483,
"pieces": 2,
"piece_tokens": [
15815,
5730
]
},
{
"object": [
"ZCL_DEMO_ABAP_DATE_TIME (CLAS)"
],
"tokens": 18634,
"pieces": 2,
"piece_tokens": [
15577,
3138
]
},
{
"object": [
"ZCL_DEMO_ABAP_XML_JSON (CLAS)"
],
"tokens": 16825,
"pieces": 2,
"piece_tokens": [
15753,
1154
]
},
{
"object": [
"ZCL_DEMO_ABAP_SQL (CLAS)"
],
"tokens": 18283,
"pieces": 2,
"piece_tokens": [
12526,
5831
]
},
{
"object": [
"ZCL_DEMO_ABAP_INTERNAL_TABLES__V755 (CLAS)"
],
"tokens": 17309,
"pieces": 2,
"piece_tokens": [
2124,
15283
]
},
{
"object": [
"ZCL_DEMO_ABAP_DYNAMIC_PROG__V756 (CLAS)"
],
"tokens": 20979,
"pieces": 2,
"piece_tokens": [
15831,
5210
]
},
{
"object": [
"16_Data_Types_and_Objects__P02 (DOC)"
],
"tokens": 17891,
"pieces": 2,
"piece_tokens": [
14421,
3541
]
}
],
"iterations_2_epochs": 748
} }

View File

@@ -1,12 +1,18 @@
"""Stage 1 data preparation: real token counts, length filter, split by document, report. """Stage 1 data preparation (Kral decisions 2026-10-03).
train/.venv/bin/python train/prepare.py [--corpus PATH] [--max-len 16384] [--seed 20261003] train/.venv/bin/python train/prepare.py [--corpus PATH] [--max-len 16384] [--seed 20261003]
Output: train/data/train.jsonl, valid.jsonl, test.jsonl (copy of valid: `mlx_lm.lora --test` needs a test file), 1. Real token counts (tokenizer of the base model).
train/data/report.md, train/data/stats.json. Only the tokenizer is loaded (no model). 2. Version dedup: in each family keep the newest version; keep an older version only if its source differs
from the newest by more than 5 % of lines (difflib).
3. Documents longer than max_len are split with the real tokenizer: classes at ENDMETHOD boundaries, markdown
at "##" headings (fallbacks: "###", then lines). Header lines are repeated in each piece.
4. Split 95/5 by family (all versions and pieces of one family stay in one split).
Output: train/data/{train,valid,test}.jsonl (test = copy of valid), report.md, stats.json. Only the tokenizer is loaded.
""" """
import argparse import argparse
import collections import collections
import difflib
import json import json
import os import os
import random import random
@@ -16,12 +22,155 @@ from transformers import AutoTokenizer
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
OUT = os.path.join(ROOT, "train", "data") OUT = os.path.join(ROOT, "train", "data")
VER = re.compile(r"__V\d+")
FAM = re.compile(r"__(V\d+|P\d+)")
RANK = {"main": 100, "v816": 90, "v758": 80, "v757": 70, "v756": 60, "v755": 50}
MARK = re.compile(r"^\* ---- .* ----$")
def pct(v, p): def pct(v, p):
return v[min(len(v) - 1, int(len(v) * p))] return v[min(len(v) - 1, int(len(v) * p))]
def version_of(d):
parts = d["package_path"].split("/")
return parts[1] if len(parts) > 1 else parts[0]
def norm_lines(text):
out = []
for i, ln in enumerate(text.split("\n")):
if i == 0 or MARK.match(ln):
continue
ln = VER.sub("", ln).rstrip()
if ln.strip():
out.append(ln)
return out
def diff_fraction(a, b):
n = max(len(a), len(b))
if n == 0:
return 0.0
m = sum(x.size for x in difflib.SequenceMatcher(None, a, b, autojunk=False).get_matching_blocks())
return (n - m) / n
# ---------------------------------------------------------------- splitting
def split_blocks(text, kind):
"""Cut the body into atomic blocks: after each ENDMETHOD. (code), before each '## ' heading (markdown)."""
lines = text.split("\n")
blocks, cur = [], []
for ln in lines:
if kind == "md" and ln.startswith("## ") and cur:
blocks.append("\n".join(cur))
cur = []
cur.append(ln)
if kind == "code" and re.match(r"^\s*ENDMETHOD\.", ln, re.I):
blocks.append("\n".join(cur))
cur = []
if cur:
blocks.append("\n".join(cur))
return blocks
def split_document(d, tok, max_len):
"""Return a list of piece texts (each <= max_len tokens, as far as possible) and the number of hard splits."""
text = d["text"]
ntok = lambda s: len(tok(s, add_special_tokens=False)["input_ids"])
kind = "md" if d["types"] == ["DOC"] else "code"
lines = text.split("\n")
head = lines[0]
# sections: marker line starts a section; body until next marker
secs, cur_marker, cur = [], None, []
for ln in lines[1:]:
if MARK.match(ln):
secs.append((cur_marker, cur))
cur_marker, cur = ln, []
else:
cur.append(ln)
secs.append((cur_marker, cur))
units = [] # (marker, block_text)
for marker, body in secs:
body_text = "\n".join(body).strip("\n")
if not body_text:
continue
sub = split_blocks(body_text, kind)
for b in sub:
units.append((marker, b))
hard = 0
# oversize unit: fall back (### for markdown), then line cut
fixed = []
for marker, b in units:
if ntok(b) <= max_len - 400:
fixed.append((marker, b))
continue
pieces = []
if kind == "md":
cur_p = []
for ln in b.split("\n"):
if ln.startswith("### ") and cur_p:
pieces.append("\n".join(cur_p))
cur_p = []
cur_p.append(ln)
pieces.append("\n".join(cur_p))
else:
pieces = [b]
for p in pieces:
if ntok(p) <= max_len - 400:
fixed.append((marker, p))
continue
hard += 1
ls, buf = p.split("\n"), []
for ln in ls:
buf.append(ln)
if ntok("\n".join(buf)) > max_len - 600:
fixed.append((marker, "\n".join(buf[:-1])))
buf = [ln]
fixed.append((marker, "\n".join(buf)))
# greedy packing; header repeated; state of the open implementation class tracked for code
title = next((ln for ln in lines if ln.startswith("# ")), "") if kind == "md" else ""
groups, cur, cur_tok, last_marker = [], [], 0, None
for marker, b in fixed:
t = ntok(b) + 20
if cur and cur_tok + t > max_len - 300:
groups.append(cur)
cur, cur_tok = [], 0
cur.append((marker, b))
cur_tok += t
groups.append(cur)
n = len(groups)
out = []
impl = None # last 'CLASS x IMPLEMENTATION.' still open at the start of the next piece
for gi, g in enumerate(groups):
h = head if n == 1 else f"{head} [part {gi + 1}/{n}]"
txt = [h, ""]
if gi > 0 and title and not g[0][1].startswith("# "):
pass
prev_marker = None
for marker, b in g:
if marker and marker != prev_marker:
txt.append(marker)
prev_marker = marker
if gi > 0 and b is g[0][1] and impl and kind == "code" and not re.match(r"^\s*CLASS\b", b, re.I):
txt.append(impl)
elif gi > 0 and b is g[0][1] and kind == "md" and title and b.splitlines()[0] != title:
txt.append(title)
txt.append("")
txt.append(b)
for ln in b.split("\n"):
m = re.match(r"^\s*CLASS\s+\S+\s+IMPLEMENTATION\.", ln, re.I)
if m:
impl = ln.strip()
if re.match(r"^\s*ENDCLASS\.", ln, re.I):
impl = None
if kind == "code" and impl and gi < n - 1:
txt.append("ENDCLASS.")
out.append("\n".join(txt).rstrip("\n") + "\n")
return out, hard
def main(): def main():
ap = argparse.ArgumentParser() ap = argparse.ArgumentParser()
ap.add_argument("--corpus", default=os.path.expanduser("~/projects/abap-llm/corpus/corpus.jsonl")) ap.add_argument("--corpus", default=os.path.expanduser("~/projects/abap-llm/corpus/corpus.jsonl"))
@@ -29,23 +178,78 @@ def main():
ap.add_argument("--max-len", type=int, default=16384) ap.add_argument("--max-len", type=int, default=16384)
ap.add_argument("--seed", type=int, default=20261003) ap.add_argument("--seed", type=int, default=20261003)
ap.add_argument("--valid-frac", type=float, default=0.05) ap.add_argument("--valid-frac", type=float, default=0.05)
ap.add_argument("--diff", type=float, default=0.05)
a = ap.parse_args() a = ap.parse_args()
tok = AutoTokenizer.from_pretrained(a.model) tok = AutoTokenizer.from_pretrained(a.model)
ntok = lambda s: len(tok(s, add_special_tokens=False)["input_ids"])
docs = [json.loads(l) for l in open(a.corpus)] docs = [json.loads(l) for l in open(a.corpus)]
for i, d in enumerate(docs): for i, d in enumerate(docs):
d["_id"] = i d["_id"] = i
d["real_tokens"] = len(tok(d["text"], add_special_tokens=False)["input_ids"]) d["real_tokens"] = ntok(d["text"])
est = sum(d["tokens"] for d in docs) d["fam"] = VER.sub("", d["objects"][0]) if d["objects"] else str(i)
real = sum(d["real_tokens"] for d in docs) d["ver"] = version_of(d)
kept = [d for d in docs if d["real_tokens"] <= a.max_len] est, real0 = sum(d["tokens"] for d in docs), sum(d["real_tokens"] for d in docs)
removed = [d for d in docs if d["real_tokens"] > a.max_len]
# Family = same demo in several release versions (__V755 ...) or parts of one text (__P01 ...). A family # ---- version dedup (an object version = all records of one object in one branch)
# is never split between train and valid (near-duplicates would make the valid loss too good). objv = collections.defaultdict(list)
for d in docs:
objv[(d["fam"], d["ver"])].append(d)
byfam = collections.defaultdict(dict)
for (f, v), ds in objv.items():
byfam[f][v] = ds
removed_versions, kept_versions, strict_extra = [], 0, 0
keep_ids = set()
n_multi = 0
for f, vers in byfam.items():
order = sorted(vers, key=lambda v: -RANK.get(v, 40))
newest = order[0]
for d in vers[newest]:
keep_ids.add(d["_id"])
kept_versions += 1
if len(order) > 1:
n_multi += 1
base = [x for d in vers[newest] for x in norm_lines(d["text"])]
kept_lines = [base]
for v in order[1:]:
lines = [x for d in vers[v] for x in norm_lines(d["text"])]
frac = diff_fraction(base, lines)
if frac > a.diff:
kept_versions += 1
for d in vers[v]:
keep_ids.add(d["_id"])
# info only: would a stricter rule (compare with all kept versions) remove it?
if min(diff_fraction(k, lines) for k in kept_lines) <= a.diff:
strict_extra += 1
kept_lines.append(lines)
else:
removed_versions.append({"object": f, "version": v, "newest": newest, "diff": round(frac, 3),
"tokens": sum(d["real_tokens"] for d in vers[v])})
after_dedup = [d for d in docs if d["_id"] in keep_ids]
real1 = sum(d["real_tokens"] for d in after_dedup)
# ---- split long documents
items, split_info, hard_total = [], [], 0
for d in after_dedup:
if d["real_tokens"] <= a.max_len:
items.append({"text": d["text"], "types": d["types"], "fam": FAM.sub("", d["objects"][0]) if d["objects"] else str(d["_id"]),
"src": d["objects"][:1], "tokens": d["real_tokens"], "pieces": 1})
continue
pieces, hard = split_document(d, tok, a.max_len)
hard_total += hard
for p in pieces:
items.append({"text": p, "types": d["types"], "fam": FAM.sub("", d["objects"][0]) if d["objects"] else str(d["_id"]),
"src": d["objects"][:1], "tokens": ntok(p), "pieces": len(pieces)})
split_info.append({"object": d["objects"][:1], "tokens": d["real_tokens"], "pieces": len(pieces),
"piece_tokens": [x["tokens"] for x in items[-len(pieces):]]})
over = [it for it in items if it["tokens"] > a.max_len]
kept = [it for it in items if it["tokens"] <= a.max_len]
real2 = sum(it["tokens"] for it in items)
# ---- split by family
fam = collections.defaultdict(list) fam = collections.defaultdict(list)
for d in kept: for it in kept:
fam[re.sub(r"__(V\d+|P\d+)", "", d["objects"][0] if d["objects"] else str(d["_id"]))].append(d) fam[it["fam"]].append(it)
keys = sorted(fam) keys = sorted(fam)
random.Random(a.seed).shuffle(keys) random.Random(a.seed).shuffle(keys)
n_valid, valid, train = max(1, round(len(kept) * a.valid_frac)), [], [] n_valid, valid, train = max(1, round(len(kept) * a.valid_frac)), [], []
@@ -55,58 +259,71 @@ def main():
os.makedirs(OUT, exist_ok=True) os.makedirs(OUT, exist_ok=True)
for name, part in (("train", train), ("valid", valid), ("test", valid)): for name, part in (("train", train), ("valid", valid), ("test", valid)):
with open(os.path.join(OUT, name + ".jsonl"), "w") as f: with open(os.path.join(OUT, name + ".jsonl"), "w") as f:
for d in part: for it in part:
f.write(json.dumps({"text": d["text"]}, ensure_ascii=False) + "\n") f.write(json.dumps({"text": it["text"]}, ensure_ascii=False) + "\n")
def summ(ds): def summ(ds):
v = sorted(d["real_tokens"] for d in ds) v = sorted(x["tokens"] for x in ds)
return {"docs": len(ds), "tokens": sum(v), "min": v[0], "p50": pct(v, .5), "p90": pct(v, .9), return {"docs": len(ds), "tokens": sum(v), "min": v[0], "p50": pct(v, .5), "p90": pct(v, .9),
"p99": pct(v, .99), "max": v[-1]} if v else {"docs": 0, "tokens": 0} "p99": pct(v, .99), "max": v[-1]}
def share(ds):
c, tot = collections.Counter(), sum(x["tokens"] for x in ds)
for x in ds:
c["DOC" if x["types"] == ["DOC"] else "CLAS" if x["types"] == ["CLAS"] else "other"] += x["tokens"]
return {k: [v, round(100 * v / tot, 1)] for k, v in c.items()}
by_type = collections.defaultdict(lambda: [0, 0]) by_type = collections.defaultdict(lambda: [0, 0])
for d in kept: for it in kept:
k = "+".join(d["types"]) if len(d["types"]) <= 2 else "mixed(%d)" % len(d["types"]) k = "+".join(it["types"]) if len(it["types"]) <= 2 else "mixed(%d)" % len(it["types"])
by_type[k][0] += 1 by_type[k][0] += 1
by_type[k][1] += d["real_tokens"] by_type[k][1] += it["tokens"]
bins = [512, 1024, 2048, 4096, 8192, 16384, 10 ** 9]
hist = collections.Counter()
for d in docs:
for b in bins:
if d["real_tokens"] <= b:
hist[b] += 1
break
stats = {"corpus": a.corpus, "seed": a.seed, "max_len": a.max_len, "estimate_tokens": est, stats = {"corpus": a.corpus, "seed": a.seed, "max_len": a.max_len, "estimate_tokens": est,
"real_tokens": real, "ratio_real_to_estimate": round(real / est, 3), "all": summ(docs), "records": len(docs), "tokens_before_dedup": real0, "tokens_after_dedup": real1,
"kept": summ(kept), "train": summ(train), "valid": summ(valid), "records_after_dedup": len(after_dedup), "tokens_after_split": real2,
"removed": [{"objects": d["objects"][:3], "tokens": d["real_tokens"]} for d in removed], "objects_with_versions": n_multi, "families": len(byfam), "versions_kept": kept_versions,
"by_type": {k: {"docs": v[0], "tokens": v[1]} for k, v in by_type.items()}, "versions_removed": len(removed_versions), "strict_rule_would_remove_more": strict_extra,
"source": sorted({d.get("source", "?") for d in docs}), "removed_versions": removed_versions, "split_docs": len(split_info), "hard_line_splits": hard_total,
"iterations_2_epochs": len(train) * 2} "pieces_over_limit": len(over), "train": summ(train), "valid": summ(valid), "all": summ(kept),
"share_before_dedup": share([{"types": d["types"], "tokens": d["real_tokens"]} for d in docs]),
"share_final": share(kept), "by_type": {k: {"docs": v[0], "tokens": v[1]} for k, v in by_type.items()},
"split_info": split_info, "iterations_2_epochs": len(train) * 2}
json.dump(stats, open(os.path.join(OUT, "stats.json"), "w"), indent=1) json.dump(stats, open(os.path.join(OUT, "stats.json"), "w"), indent=1)
L = ["# Stage 1 data report", "", L = ["# Stage 1 data report", "",
f"Corpus: `{a.corpus}` (source {', '.join(stats['source'])}, Apache-2.0, see corpus/out/ATTRIBUTION.md).", f"Corpus: `{a.corpus}` (SAP-samples/abap-cheat-sheets, Apache-2.0, see corpus/out/ATTRIBUTION.md).",
f"Tokenizer: `{a.model}`. Seed {a.seed}. max_seq_length {a.max_len}.", f"Tokenizer: `{a.model}`. Seed {a.seed}. max_seq_length {a.max_len}.", "",
f"Split by family (release versions and parts of one document stay together): {len(fam)} families.", "", "## Version dedup (keep the newest; keep an older version only if it differs by more than "
"| | docs | tokens |", "|---|---|---|", f"{int(a.diff * 100)} % of lines)", "",
f"| corpus | {len(docs)} | {real} (estimate chars/4: {est}, ratio {stats['ratio_real_to_estimate']}) |", f"- Object families: {len(byfam)}; with more than one version: {n_multi}.",
f"| removed (> {a.max_len}) | {len(removed)} | {sum(d['real_tokens'] for d in removed)} |", f"- Versions kept: {kept_versions}. Versions removed: {len(removed_versions)}.",
f"- Records: {len(docs)} before, {len(after_dedup)} after.",
f"- Tokens before: {real0} (chars/4 estimate {est}). After dedup: {real1}. After splitting: {real2}.",
f"- Info: a stricter rule (compare also with the other kept versions) would remove {strict_extra} more kept versions.",
"", "## Splitting of documents over the limit", "",
f"- Documents split: {len(split_info)}; pieces: {sum(s['pieces'] for s in split_info)}; hard line splits "
f"(a block was too big): {hard_total}; pieces still over the limit: {len(over)}.", "",
"## Result", "", "| | docs | tokens |", "|---|---|---|",
f"| train | {len(train)} | {stats['train']['tokens']} |", f"| train | {len(train)} | {stats['train']['tokens']} |",
f"| valid (= test file) | {len(valid)} | {stats['valid']['tokens']} |", "", f"| valid (= test file) | {len(valid)} | {stats['valid']['tokens']} |", "",
f"Iterations for 2 epochs (batch 1): {len(train) * 2}.", "", f"Split by family: {len(fam)} families. **Iterations for 2 epochs (batch 1): {len(train) * 2}.**", "",
"## Token distribution per document (real tokenizer)", "", "| set | min | p50 | p90 | p99 | max |", "|---|---|---|---|---|---|"] "## Token share by type", "", "| | before dedup | final (kept pieces) |", "|---|---|---|"]
for n in ("all", "kept", "train", "valid"): for k in ("DOC", "CLAS", "other"):
s = stats[n] b, f = stats["share_before_dedup"].get(k, [0, 0]), stats["share_final"].get(k, [0, 0])
L.append(f"| {n} | {s['min']} | {s['p50']} | {s['p90']} | {s['p99']} | {s['max']} |") L.append(f"| {k} | {b[0]} ({b[1]} %) | {f[0]} ({f[1]} %) |")
L += ["", "| tokens per document up to | docs |", "|---|---|"] L += ["", "## Token distribution per document (final)", "", "| set | min | p50 | p90 | p99 | max |", "|---|---|---|---|---|---|"]
L += [f"| {b if b < 10 ** 9 else 'more'} | {hist[b]} |" for b in bins] for n_ in ("all", "train", "valid"):
L += ["", "## By object type (kept documents)", "", "| types | docs | tokens |", "|---|---|---|"] s = stats[n_]
L.append(f"| {n_} | {s['min']} | {s['p50']} | {s['p90']} | {s['p99']} | {s['max']} |")
L += ["", "## By object type (final)", "", "| types | docs | tokens |", "|---|---|---|"]
for k, v in sorted(by_type.items(), key=lambda kv: -kv[1][1]): for k, v in sorted(by_type.items(), key=lambda kv: -kv[1][1]):
L.append(f"| {k} | {v[0]} | {v[1]} |") L.append(f"| {k} | {v[0]} | {v[1]} |")
L += ["", "## Removed documents", ""] L += ["", "## Removed versions", "", "| object | version | newest | diff | tokens |", "|---|---|---|---|---|"]
L += [f"- {', '.join(d['objects'][:2])}: {d['real_tokens']} tokens" for d in removed] or ["None."] L += [f"| {r['object']} | {r['version']} | {r['newest']} | {r['diff']} | {r['tokens']} |" for r in removed_versions]
L += ["", "## Split documents", ""]
L += [f"- {s['object']}: {s['tokens']} tokens -> {s['pieces']} pieces {s['piece_tokens']}" for s in split_info]
open(os.path.join(OUT, "report.md"), "w").write("\n".join(L) + "\n") open(os.path.join(OUT, "report.md"), "w").write("\n".join(L) + "\n")
print("\n".join(L)) print("\n".join(L[:60]))
if __name__ == "__main__": if __name__ == "__main__":