Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# Stage 1 state
|
||||
|
||||
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:30.
|
||||
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 23:00.
|
||||
|
||||
## Done
|
||||
|
||||
@@ -15,14 +15,16 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
|
||||
run directories `runs/stage1/<label>/`).
|
||||
|
||||
- Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): `train/prepare.py`, report `train/data/report.md`.
|
||||
Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
|
||||
Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept;
|
||||
tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces
|
||||
(classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece
|
||||
over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %.
|
||||
Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
|
||||
**618 -> 748 iterations for 2 epochs.** `test.jsonl` = copy of valid.
|
||||
- Step 1 done (2026-10-03 23:00, Kral decisions applied, strict dedup rule): `train/prepare.py`, report
|
||||
`train/data/report.md`. Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
|
||||
Version dedup: main (ABAP Cloud) and the newest v* (Standard ABAP) always kept; each older v* is compared with
|
||||
the next newer v* and kept only if more than 5 % of lines differ. 335 object versions: 323 kept, 12 removed.
|
||||
Tokens 2,806,501 -> 2,658,082 after dedup. Documents over 16384 are split, not removed: 41 documents -> 82
|
||||
pieces (classes at ENDMETHOD, markdown at "##"; 15 blocks needed a line cut); tokens after split 2,660,834;
|
||||
no piece over the limit. Token share: DOC 49.0 %, CLAS 46.9 %, other 4.2 %.
|
||||
Train 374 docs / 2.53M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
|
||||
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
|
||||
`test.jsonl` = copy of valid.
|
||||
|
||||
## Running (detached)
|
||||
|
||||
@@ -34,6 +36,8 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
|
||||
## Next
|
||||
|
||||
0. After the baseline ends: Step 3 training test (20 iterations), Kral stops A4H first and the MLX server is stopped.
|
||||
|
||||
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
||||
(`t01_test_budget40` holds the T01 test result). Commit.
|
||||
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
|
||||
|
||||
Reference in New Issue
Block a user