Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 21:58:04 +02:00
parent f0593933f9
commit e083c9ca13
6 changed files with 593 additions and 5 deletions

View File

@@ -15,6 +15,11 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
run directories `runs/stage1/<label>/`).
- Step 1 done (2026-10-03 22:00): `train/prepare.py`, report `train/data/report.md`. Corpus replaced by
SAP-samples/abap-cheat-sheets (Apache-2.0): 370 docs, 2.81M real tokens. 45 docs (840k tokens, 30 %) are longer
than 16384 and removed. Train 309 docs / 1.89M tokens, valid 16 docs / 80k tokens (split by family, seed
20261003), 618 iterations for 2 epochs. `test.jsonl` = copy of valid (needed by `mlx_lm.lora --test`).
## Running (detached)
- MLX server PID 59352, log `runs/stage1/server.log`.
@@ -27,8 +32,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Step 1: `train/prepare.py` (real token counts with the base model tokenizer, length filter 16384,
95/5 split by document, `train/data/train.jsonl` and `valid.jsonl`, report). A4H is not needed.
2. Decisions for Kral on the data (see report): the 45 removed docs, DOC share, valid size.
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20