92 lines
7.1 KiB
Python
92 lines
7.1 KiB
Python
"""Writes docs/stage2-build.md from runs/stage2_data/build_report.json (run after train/build_stage2.py)."""
|
|
import json
|
|
import os
|
|
|
|
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
|
rep = json.load(open(os.path.join(ROOT, "runs", "stage2_data", "build_report.json")))
|
|
tr, va, rs = rep["train"], rep["valid"], rep["reserve"]
|
|
s1 = rep["stage1_stats"]["train"]["tokens"]
|
|
s1d = rep["stage1_stats"]["train"]["docs"]
|
|
L2 = tr["loss_tokens"]
|
|
|
|
|
|
def row(e1, e2, L=L2):
|
|
a, b = s1 * e1, L * e2
|
|
return f"| {e1} | {e2} | {a / 1e6:.2f} M | {b / 1e6:.2f} M | {100 * b / (a + b):.0f} % | {(a + b) / 1e6:.1f} M |"
|
|
|
|
|
|
drop = rep["dropped"]
|
|
dropped_lines = "\n".join(f"- {k}: {v}" for k, v in drop.items())
|
|
kinds = "\n".join(f"| {k} | {v['samples']} | {v['families']} | {v['short']} |" for k, v in rep["kinds"].items())
|
|
doc = f"""# Stage 2 training set, build of {rep['built']}
|
|
|
|
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
|
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
|
|
|
## Result
|
|
{rep['steps']['accepted_trajectories']} accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> {rep['steps']['after_eval_overlap']} after the eval overlap check
|
|
-> {rep['steps']['after_token_limit']} after the 48k limit (none cut) -> CLAS cap 35 %: {rep['steps']['clas_cap']['clas_kept']} CLAS kept, **{rep['steps']['clas_cap']['clas_reserve']} CLAS in reserve** (`stage2_reserve.jsonl`)
|
|
-> **{tr['samples']} train + {va['samples']} valid samples** ({tr['tokens'] / 1e6:.2f} M + {va['tokens'] / 1e6:.2f} M tokens, {tr['loss_tokens'] / 1e6:.2f} M loss tokens in train, p50 {tr['p50']}, p95 {tr['p95']}, max {tr['max']}, repair share {tr['repair_share']}).
|
|
Dropped:
|
|
{dropped_lines}
|
|
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
|
|
|
## Not good enough yet (kinds under the minimum of {rep['settings']['min_per_kind']})
|
|
| kind | samples | families | short |
|
|
|---|---|---|---|
|
|
{kinds}
|
|
|
|
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
|
|
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
|
|
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
|
|
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
|
|
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
|
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
|
|
|
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
|
|
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
|
|
|
|
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
|
|---|---|---|---|---|---|
|
|
{row(2, 3)}
|
|
{row(1, 3)}
|
|
{row(0.5, 3)}
|
|
{row(0.33, 3)}
|
|
{row(1, 3, 3 * L2)}
|
|
|
|
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
|
|
|
|
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
|
|
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
|
|
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
|
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
|
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
|
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
|
|
|
|
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
|
|
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
|
|
|
|
| class of the trajectory | weight | reason |
|
|
|---|---|---|
|
|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
|
|
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
|
|
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
|
|
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
|
|
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
|
|
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
|
|
|
|
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
|
|
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
|
|
|
|
## Other decisions of 2026-10-06
|
|
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
|
|
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
|
|
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
|
|
- **Memory test:** waits until the data is near the size of the first SFT run.
|
|
|
|
## Also built
|
|
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
|
"""
|
|
open(os.path.join(ROOT, "docs", "stage2-build.md"), "w").write(doc)
|
|
print("written", len(doc))
|