Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -15,6 +15,11 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
- Runner: `train/baseline.py --label <label>` (one task at a time, results `runs/stage1/<label>.json`,
|
||||
run directories `runs/stage1/<label>/`).
|
||||
|
||||
- Step 1 done (2026-10-03 22:00): `train/prepare.py`, report `train/data/report.md`. Corpus replaced by
|
||||
SAP-samples/abap-cheat-sheets (Apache-2.0): 370 docs, 2.81M real tokens. 45 docs (840k tokens, 30 %) are longer
|
||||
than 16384 and removed. Train 309 docs / 1.89M tokens, valid 16 docs / 80k tokens (split by family, seed
|
||||
20261003), 618 iterations for 2 epochs. `test.jsonl` = copy of valid (needed by `mlx_lm.lora --test`).
|
||||
|
||||
## Running (detached)
|
||||
|
||||
- MLX server PID 59352, log `runs/stage1/server.log`.
|
||||
@@ -27,8 +32,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
|
||||
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
||||
(`t01_test_budget40` holds the T01 test result). Commit.
|
||||
2. Step 1: `train/prepare.py` (real token counts with the base model tokenizer, length filter 16384,
|
||||
95/5 split by document, `train/data/train.jsonl` and `valid.jsonl`, report). A4H is not needed.
|
||||
2. Decisions for Kral on the data (see report): the 45 removed docs, DOC share, valid size.
|
||||
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
||||
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
||||
|
||||
Reference in New Issue
Block a user