6.2 KiB
Stage 2 training set, build of 2026-10-06 09:03:41
Builder: train/build_stage2.py (output runs/stage2_data/, data card README.md, report build_report.json; this page: train/build_doc.py). Hook for item D: train/hooks_example.py.
HF dataset (private): erhankeseli/abap-stage2-data. The local Qwen trajectories (series A) are not read.
Result
81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, 35 CLAS in reserve (stage2_reserve.jsonl)
-> 36 train + 4 valid samples (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
Dropped:
- DDLS: {'over_48k': 5}
- INTF: {'over_48k': 1} Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
Not good enough yet (kinds under the minimum of 25)
| kind | samples | families | short |
|---|---|---|---|
| CLAS | 14 | 14 | 11 |
| INTF | 2 | 2 | 23 |
| DDLS | 9 | 9 | 16 |
| FUNC | 8 | 8 | 17 |
| PROG | 5 | 5 | 20 |
| TABL | 2 | 2 | 23 |
| STRU | 0 | 0 | 25 |
| MSAG | 0 | 0 | 25 |
| EXC | 0 | 0 | 25 |
- STRU, MSAG, exception: 0 trajectories; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects. The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- Reads of another run's object: the test system held leftover objects of earlier runs (the model's own
ZCL_<prefix>_...classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (--keep-foreign-readskeeps them). - Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
| 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M |
| 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M |
| 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M |
| 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M |
| 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M |
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. --s1-epochs takes fractions.
Decision (Kral + Opus 2026-10-06): this rule. train/hf_train_bf16.py --s2-loss-share 0.6 (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
Own-test weights (decided 2026-10-06: down-weight, do not drop; train/hooks_example.py, own_test_weight)
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
| class of the trajectory | weight | reason |
|---|---|---|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, 15 without own tests at 0.5, 2 unreliable at 0.5: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1. Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
Other decisions of 2026-10-06
- CLAS cap 35 % stays a build setting (
--clas-cap); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve). - Length limit: the memory test decides (
docs/bf16-memory.md: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then. - Foreign names: list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (
docs/foreign-objects-report.md). - Memory test: waits until the data is near the size of the first SFT run.
Also built
train/hf_train_bf16.py (bf16, loss mask, mixing, memory test), docs/bf16-memory.md (memory table by GPU, estimates), train/hooks_example.py (hook for item D, weights).