# Stage 2 training set, first build (2026-10-06) Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`). Hook for item D: `train/hooks_example.py`. HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read. ## Result 90 accepted DeepSeek trajectories -> 90 after the eval overlap check (0 removed) -> 83 after the 48k limit (**7 dropped over 48000 tokens: 6 DDLS, 1 INTF**; none cut) -> CLAS cap 35 %: 18 CLAS kept, **31 CLAS in reserve** (`stage2_reserve.jsonl`) -> **47 train + 5 valid samples** (1.28 M + 0.12 M tokens, 0.41 M loss tokens in train, p50 24771, p95 41829, max 47694). Loss mask checked on all 47 train samples: no span contains a tool result, the system turn or a user turn; loss share 32 % of the tokens; no token straddles a span boundary. ## Not good enough yet (kinds under the minimum of 25) | kind | samples | families | short | |---|---|---|---| | CLAS | 18 | 18 | 7 | | INTF | 2 | 2 | 23 | | DDLS | 13 | 13 | 12 | | FUNC | 11 | 11 | 14 | | PROG | 5 | 5 | 20 | | TABL | 3 | 3 | 22 | | STRU | 0 | 0 | 25 | | MSAG | 0 | 0 | 25 | | EXC | 0 | 0 | 25 | - **STRU, MSAG, exception: 0 trajectories**; INTF 2, TABL 3, PROG 5. Only CLAS, DDLS and FUNC are near their share. Nothing is filled with copies. This is what the restart plan (new kinds first) has to fix. - **DDLS loses 6 of 19 trajectories to the 48k limit.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (memory test decides), or generate CDS tasks with shorter runs. Not decided. - Validation: 5 samples ({'CLAS': 2, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt. - 2 trajectories of one task and K variants are on the same side (family key); there are no same-task pairs yet in practice. - Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60). ## Stage 1 : stage 2 ratio (proposal, Kral decides) Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 47 samples, 0.41 M loss tokens per epoch. Loss tokens per option (the table uses today's data): | stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* | |---|---|---|---|---|---| | 2 | 3 | 5.05 M | 1.23 M | 20 % | 6.3 M | | 1 | 3 | 2.53 M | 1.23 M | 33 % | 3.8 M | | 0.5 | 3 | 1.26 M | 1.23 M | 49 % | 2.5 M | | 0.33 | 3 | 0.83 M | 1.23 M | 60 % | 2.1 M | | 1 | 3 | 2.53 M | 3.70 M | 59 % | 6.2 M | (*stage 1 tokens plus all stage 2 tokens per epoch; last row: if stage 2 had three times today's data, as expected after the restart.) **Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).** Reasons: (1) stage 2 is the behavior we want (repair after the first error, write after a few reads, the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against 1.2 M of stage 2: the model would mostly learn documents again, the format of the tool calls would be a small part. (3) With 47 samples, more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss (5 samples today) is too thin to catch it, so keep an eye on the train loss curve and use the checkpoints. (4) The rule scales with the data: today it means about 0.33 epochs of stage 1 (about 125 documents), with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions. Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1. ## Also built `train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).