Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 22:09:15 +02:00
parent e083c9ca13
commit 0677d035da
5 changed files with 1014 additions and 421 deletions

View File

@@ -13,8 +13,9 @@ Stage 2 (SFT on teacher trajectories) is not part of this task.
Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03):
SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`).
Real numbers from Step 1 (`train/data/report.md`): 370 documents, 2,806,501 tokens (Qwen tokenizer;
the chars/4 estimate is 2.64M). The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line:
Real numbers from Step 1 (`train/data/report.md`): 370 records, 2,806,501 tokens (Qwen tokenizer;
the chars/4 estimate is 2.64M). After version dedup and splitting: train 374 documents / 2.57M tokens,
valid 20 documents / 135k tokens, 748 iterations for 2 epochs. The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line:
`{"text", "objects", "types", "language_version", "package_path",
"grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`.