Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
# Stage 2 training set, build of 2026-10-06 07:36:30
|
||||
# Stage 2 training set, build of 2026-10-06 09:03:41
|
||||
|
||||
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
||||
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||
@@ -8,9 +8,7 @@ HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectorie
|
||||
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
|
||||
Dropped:
|
||||
- FUNC: {'read_of_another_runs_object': 3}
|
||||
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
|
||||
- TABL: {'read_of_another_runs_object': 1}
|
||||
- DDLS: {'over_48k': 5}
|
||||
- INTF: {'over_48k': 1}
|
||||
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
||||
|
||||
@@ -34,7 +32,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
|
||||
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
|
||||
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
@@ -52,7 +50,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
|
||||
|
||||
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
|
||||
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
|
||||
|
||||
| class of the trajectory | weight | reason |
|
||||
|---|---|---|
|
||||
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
|
||||
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
|
||||
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
|
||||
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
|
||||
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
|
||||
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
|
||||
|
||||
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
|
||||
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
|
||||
|
||||
## Other decisions of 2026-10-06
|
||||
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
|
||||
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
|
||||
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
|
||||
- **Memory test:** waits until the data is near the size of the first SFT run.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
|
||||
Reference in New Issue
Block a user