diff --git a/.gitignore b/.gitignore index 93004b1..fef733c 100644 --- a/.gitignore +++ b/.gitignore @@ -4,3 +4,4 @@ runs/ __pycache__/ .DS_Store train/.venv/ +train/data/*.jsonl diff --git a/docs/stage1-training-task.md b/docs/stage1-training-task.md index f622888..fa2f5f3 100644 --- a/docs/stage1-training-task.md +++ b/docs/stage1-training-task.md @@ -11,10 +11,12 @@ Read CLAUDE.md first. This task adds stage 1 of the training plan: Stage 2 (SFT on teacher trajectories) is not part of this task. -Input: `~/projects/abap-llm/corpus/corpus.jsonl` (about 1,765 documents, -about 1.4M tokens, estimate). One record per line: +Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03): +SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`). +Real numbers from Step 1 (`train/data/report.md`): 370 documents, 2,806,501 tokens (Qwen tokenizer; +the chars/4 estimate is 2.64M). The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line: `{"text", "objects", "types", "language_version", "package_path", -"grouped", "obsolete", "tokens"}`. Token counts are estimates (chars / 4). +"grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`. ## Rules diff --git a/train/STATE.md b/train/STATE.md index 17b5e8c..68d537a 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -15,6 +15,11 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U - Runner: `train/baseline.py --label