Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -11,10 +11,12 @@ Read CLAUDE.md first. This task adds stage 1 of the training plan:
|
||||
|
||||
Stage 2 (SFT on teacher trajectories) is not part of this task.
|
||||
|
||||
Input: `~/projects/abap-llm/corpus/corpus.jsonl` (about 1,765 documents,
|
||||
about 1.4M tokens, estimate). One record per line:
|
||||
Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03):
|
||||
SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`).
|
||||
Real numbers from Step 1 (`train/data/report.md`): 370 documents, 2,806,501 tokens (Qwen tokenizer;
|
||||
the chars/4 estimate is 2.64M). The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line:
|
||||
`{"text", "objects", "types", "language_version", "package_path",
|
||||
"grouped", "obsolete", "tokens"}`. Token counts are estimates (chars / 4).
|
||||
"grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`.
|
||||
|
||||
## Rules
|
||||
|
||||
|
||||
Reference in New Issue
Block a user