# Stage 2 training set, build of 2026-10-06 07:36:30 Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`. HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read. ## Result 81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check -> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`) -> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86). Dropped: - FUNC: {'read_of_another_runs_object': 3} - DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5} - TABL: {'read_of_another_runs_object': 1} - INTF: {'over_48k': 1} Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary. ## Not good enough yet (kinds under the minimum of 25) | kind | samples | families | short | |---|---|---|---| | CLAS | 14 | 14 | 11 | | INTF | 2 | 2 | 23 | | DDLS | 9 | 9 | 16 | | FUNC | 8 | 8 | 17 | | PROG | 5 | 5 | 20 | | TABL | 2 | 2 | 23 | | STRU | 0 | 0 | 25 | | MSAG | 0 | 0 | 25 | | EXC | 0 | 0 | 25 | - **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this. - **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided. - **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL__...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them). - Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt. - Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60). ## Stage 1 : stage 2 ratio (proposal, Kral decides) Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data): | stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* | |---|---|---|---|---|---| | 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M | | 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M | | 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M | | 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M | | 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M | (*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.) **Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).** Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again. (3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints. (4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions. Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1. ## Also built `train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).