Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 22:09:15 +02:00
parent e083c9ca13
commit 0677d035da
5 changed files with 1014 additions and 421 deletions

View File

@@ -1,6 +1,6 @@
# Stage 1 state
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 19:10.
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 22:30.
## Done
@@ -15,10 +15,14 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:00): `train/prepare.py`, report `train/data/report.md`. Corpus replaced by
SAP-samples/abap-cheat-sheets (Apache-2.0): 370 docs, 2.81M real tokens. 45 docs (840k tokens, 30 %) are longer
than 16384 and removed. Train 309 docs / 1.89M tokens, valid 16 docs / 80k tokens (split by family, seed
20261003), 618 iterations for 2 epochs. `test.jsonl` = copy of valid (needed by `mlx_lm.lora --test`).
- Step 1 done (2026-10-03 22:30, Kral decisions 1-4 applied): `train/prepare.py`, report `train/data/report.md`.
Corpus: SAP-samples/abap-cheat-sheets (Apache-2.0), 370 records, 2.81M real tokens.
Version dedup (newest kept; older kept only if > 5 % of lines differ): 335 object versions, 16 removed, 319 kept;
tokens 2,806,501 -> 2,699,110. Documents over 16384 are split, not removed: 43 documents -> 86 pieces
(classes at ENDMETHOD, markdown at "##"; 17 blocks needed a line cut); tokens after split 2,701,996; no piece
over the limit. Token share after the changes: DOC 48.2 %, CLAS 48.0 %, other 3.8 %.
Train 374 docs / 2.57M tokens, valid 20 docs / 135k tokens (split by family, seed 20261003).
**618 -> 748 iterations for 2 epochs.** `test.jsonl` = copy of valid.
## Running (detached)
@@ -32,7 +36,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Decisions for Kral on the data (see report): the 45 removed docs, DOC share, valid size.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20