C+E: stage 2 builder (mask, 48k, CLAS cap, family split, hook), private HF dataset, bf16 mixed training script with memory test, memory table, ratio proposal
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
48
docs/bf16-memory.md
Normal file
48
docs/bf16-memory.md
Normal file
@@ -0,0 +1,48 @@
|
||||
# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
|
||||
|
||||
**No GPU job is started without Kral's go.** Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000` is the memory test, no data needed, 3 optimizer steps per length, stops at the first OOM, uploads the result as `memtest_*.json` to the output repo).
|
||||
|
||||
## Model (from config.json of Qwen3.8-27B)
|
||||
64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**.
|
||||
|
||||
## Memory by component at 48k tokens (GB; the estimate, not a measurement)
|
||||
| component | GB | how |
|
||||
|---|---|---|
|
||||
| weights bf16 | 51.8 | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | 1.0 | 107 M x 10 bytes |
|
||||
| layer inputs for the backward pass (checkpoints) | 29.3 on the GPU, about 0 with the Unsloth CPU offload | 48k x 5120 x 2 bytes x 64 layers; the offload needs 29 GB of host RAM (all flavors have 142 GB or more) |
|
||||
| recompute peak of one layer | 9.2 | MLP tensors 48k x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | 3 chunked, 67 with full logits | full logits: 48k x 248320 x (bf16 + fp32 upcast + grad); only a chunked or fused cross entropy is realistic |
|
||||
|
||||
## Fit by GPU (92 % of the card counted as usable)
|
||||
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (needs model parallel) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits |
|
||||
| 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits |
|
||||
| 16k | on the GPU | chunked | 71 | fits | fits | fits | fits |
|
||||
| 16k | on the GPU | full logits | 90 | no | tight | fits | fits |
|
||||
| 32k | CPU offload (Unsloth) | chunked | 63 | fits | fits | fits | fits |
|
||||
| 32k | CPU offload (Unsloth) | full logits | 104 | no | no | fits | fits |
|
||||
| 32k | on the GPU | chunked | 82 | no | fits | fits | fits |
|
||||
| 32k | on the GPU | full logits | 124 | no | no | fits | fits |
|
||||
| 48k | CPU offload (Unsloth) | chunked | 65 | fits | fits | fits | fits |
|
||||
| 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits |
|
||||
| 48k | on the GPU | chunked | 94 | no | tight | fits | fits |
|
||||
| 48k | on the GPU | full logits | 158 | no | no | no | fits |
|
||||
|
||||
Reading:
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, with the CPU offload) fits 32k and 48k; the A100 and the RTX PRO 6000 do not. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is not known). The memory test shows it at once: if `seq_len 16000` already fails on an H200, the loss is the cause.
|
||||
Fallback (not built yet, small): compute the hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour) is only possible with the CPU offload and a chunked loss** (about 65 GB estimated at 48k: no margin). H200 141 GB (5 USD/hour) fits with margin;
|
||||
RTX PRO 6000 96 GB (2.75 USD/hour) fits with the offload and the chunked loss (est. 65 GB).
|
||||
- Samples over 32k are a minority (p90 39560 tokens in the current data, p50 24.8k): if 48k does not fit, 32k would drop 15 of 52 samples (29 %; 10 of them DDLS), so test 48k first.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): possible only with model parallelism (`device_map`), one card works at a time, Unsloth multi-GPU is limited; not recommended before the single card test.
|
||||
|
||||
## Proposed order
|
||||
1. Memory test on **h200** (about 30 minutes, 2.5 USD): sweep 16000,32000,48000 with the offload. If 48k fits with margin: stay on H200 or try rtx-pro-6000 for the real run.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours of work), repeat the test.
|
||||
3. Real run after the stage 1 : stage 2 ratio decision (below).
|
||||
|
||||
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 is faster, H200 about 2 to 2.5 times an A100)
|
||||
Tokens of the proposed run (1 epoch stage 1 and 3 epochs of the 47 stage 2 samples as built today: 2.53 M + 3 x 1.28 M = 6.4 M tokens): A100 about 5.9 h (about 15 USD), H200 about 2.7 h (about 14 USD). The real number comes from the memory test (`step_seconds`).
|
||||
52
docs/stage2-build.md
Normal file
52
docs/stage2-build.md
Normal file
@@ -0,0 +1,52 @@
|
||||
# Stage 2 training set, first build (2026-10-06)
|
||||
|
||||
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`). Hook for item D: `train/hooks_example.py`.
|
||||
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||
|
||||
## Result
|
||||
90 accepted DeepSeek trajectories -> 90 after the eval overlap check (0 removed) -> 83 after the 48k limit
|
||||
(**7 dropped over 48000 tokens: 6 DDLS, 1 INTF**; none cut) -> CLAS cap 35 %: 18 CLAS kept, **31 CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||
-> **47 train + 5 valid samples** (1.28 M + 0.12 M tokens, 0.41 M loss tokens in train, p50 24771, p95 41829, max 47694).
|
||||
Loss mask checked on all 47 train samples: no span contains a tool result, the system turn or a user turn; loss share 32 % of the tokens; no token straddles a span boundary.
|
||||
|
||||
## Not good enough yet (kinds under the minimum of 25)
|
||||
| kind | samples | families | short |
|
||||
|---|---|---|---|
|
||||
| CLAS | 18 | 18 | 7 |
|
||||
| INTF | 2 | 2 | 23 |
|
||||
| DDLS | 13 | 13 | 12 |
|
||||
| FUNC | 11 | 11 | 14 |
|
||||
| PROG | 5 | 5 | 20 |
|
||||
| TABL | 3 | 3 | 22 |
|
||||
| STRU | 0 | 0 | 25 |
|
||||
| MSAG | 0 | 0 | 25 |
|
||||
| EXC | 0 | 0 | 25 |
|
||||
|
||||
- **STRU, MSAG, exception: 0 trajectories**; INTF 2, TABL 3, PROG 5. Only CLAS, DDLS and FUNC are near their share. Nothing is filled with copies. This is what the restart plan (new kinds first) has to fix.
|
||||
- **DDLS loses 6 of 19 trajectories to the 48k limit.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (memory test decides), or generate CDS tasks with shorter runs. Not decided.
|
||||
- Validation: 5 samples ({'CLAS': 2, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- 2 trajectories of one task and K variants are on the same side (family key); there are no same-task pairs yet in practice.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 47 samples, 0.41 M loss tokens per epoch. Loss tokens per option (the table uses today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
|---|---|---|---|---|---|
|
||||
| 2 | 3 | 5.05 M | 1.23 M | 20 % | 6.3 M |
|
||||
| 1 | 3 | 2.53 M | 1.23 M | 33 % | 3.8 M |
|
||||
| 0.5 | 3 | 1.26 M | 1.23 M | 49 % | 2.5 M |
|
||||
| 0.33 | 3 | 0.83 M | 1.23 M | 60 % | 2.1 M |
|
||||
| 1 | 3 | 2.53 M | 3.70 M | 59 % | 6.2 M |
|
||||
|
||||
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: if stage 2 had three times today's data, as expected after the restart.)
|
||||
|
||||
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
|
||||
Reasons: (1) stage 2 is the behavior we want (repair after the first error, write after a few reads, the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against 1.2 M of stage 2: the model would mostly learn documents again, the format of the tool calls would be a small part.
|
||||
(3) With 47 samples, more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss (5 samples today) is too thin to catch it, so keep an eye on the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: today it means about 0.33 epochs of stage 1 (about 125 documents), with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
Reference in New Issue
Block a user