4.5 KiB
Stage 2 training set, build of 2026-10-06 07:36:30
Builder: train/build_stage2.py (output runs/stage2_data/, data card README.md, report build_report.json; this page: train/build_doc.py). Hook for item D: train/hooks_example.py.
HF dataset (private): erhankeseli/abap-stage2-data. The local Qwen trajectories (series A) are not read.
Result
81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, 35 CLAS in reserve (stage2_reserve.jsonl)
-> 36 train + 4 valid samples (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
Dropped:
- FUNC: {'read_of_another_runs_object': 3}
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
- TABL: {'read_of_another_runs_object': 1}
- INTF: {'over_48k': 1} Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
Not good enough yet (kinds under the minimum of 25)
| kind | samples | families | short |
|---|---|---|---|
| CLAS | 14 | 14 | 11 |
| INTF | 2 | 2 | 23 |
| DDLS | 9 | 9 | 16 |
| FUNC | 8 | 8 | 17 |
| PROG | 5 | 5 | 20 |
| TABL | 2 | 2 | 23 |
| STRU | 0 | 0 | 25 |
| MSAG | 0 | 0 | 25 |
| EXC | 0 | 0 | 25 |
- STRU, MSAG, exception: 0 trajectories; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects. The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- Reads of another run's object: the test system held leftover objects of earlier runs (the model's own
ZCL_<prefix>_...classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (--keep-foreign-readskeeps them). - Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
Stage 1 : stage 2 ratio (proposal, Kral decides)
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
| 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M |
| 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M |
| 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M |
| 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M |
| 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M |
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. --s1-epochs takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
Also built
train/hf_train_bf16.py (bf16, loss mask, mixing, memory test), docs/bf16-memory.md (memory table by GPU, estimates), train/hooks_example.py (hook for item D, weights).