Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 22:23:04 +02:00
parent 0677d035da
commit 7b8ca01bde
5 changed files with 179 additions and 211 deletions

View File

@@ -1,6 +1,6 @@
# Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:30.
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 23:00.
## Done
@@ -15,14 +15,16 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): `train/prepare.py`, report `train/data/report.md`.
Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept;
tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces
(classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece
over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %.
Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**618 -> 748 iterations for 2 epochs.** `test.jsonl` = copy of valid.
- Step 1 done (2026-10-03 23:00, Kral decisions applied, strict dedup rule): `train/prepare.py`, report
`train/data/report.md`. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with
the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed.
Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82
pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834;
no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %.
Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
`test.jsonl` = copy of valid.
## Running (detached)
@@ -34,6 +36,8 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
## Next
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.