Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,21 +1,24 @@
|
||||
# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
|
||||
|
||||
**No GPU job is started without Kral's go.** Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000` is the memory test, no data needed, 3 optimizer steps per length, stops at the first OOM, uploads the result as `memtest_*.json` to the output repo).
|
||||
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
|
||||
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
|
||||
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
|
||||
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
|
||||
|
||||
## Model (from config.json of Qwen3.8-27B)
|
||||
64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**.
|
||||
|
||||
## Memory by component at 48k tokens (GB; the estimate, not a measurement)
|
||||
| component | GB | how |
|
||||
|---|---|---|
|
||||
| weights bf16 | 51.8 | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | 1.0 | 107 M x 10 bytes |
|
||||
| layer inputs for the backward pass (checkpoints) | 29.3 on the GPU, about 0 with the Unsloth CPU offload | 48k x 5120 x 2 bytes x 64 layers; the offload needs 29 GB of host RAM (all flavors have 142 GB or more) |
|
||||
| recompute peak of one layer | 9.2 | MLP tensors 48k x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | 3 chunked, 67 with full logits | full logits: 48k x 248320 x (bf16 + fp32 upcast + grad); only a chunked or fused cross entropy is realistic |
|
||||
## Memory by component (GB; estimate, not a measurement)
|
||||
| component | 48k | 64k | how |
|
||||
|---|---|---|---|
|
||||
| weights bf16 | 51.8 | 51.8 | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | 1.0 | 1.0 | 107 M x 10 bytes |
|
||||
| layer inputs for the backward pass | 29.3 on the GPU, about 0 with the Unsloth CPU offload (then 29 GB host RAM) | 39.1 / about 0 (host RAM 39 GB) | seq x 5120 x 2 bytes x 64 layers |
|
||||
| recompute peak of one layer | 9.2 | 11.3 | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | 3 chunked, 67 full logits | 3 chunked, 89 full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
|
||||
|
||||
## Fit by GPU (92 % of the card counted as usable)
|
||||
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (needs model parallel) |
|
||||
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (model parallel) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits |
|
||||
| 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits |
|
||||
@@ -29,20 +32,26 @@
|
||||
| 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits |
|
||||
| 48k | on the GPU | chunked | 94 | no | tight | fits | fits |
|
||||
| 48k | on the GPU | full logits | 158 | no | no | no | fits |
|
||||
| 64k | CPU offload (Unsloth) | chunked | 67 | fits | fits | fits | fits |
|
||||
| 64k | CPU offload (Unsloth) | full logits | 153 | no | no | no | fits |
|
||||
| 64k | on the GPU | chunked | 106 | no | no | fits | fits |
|
||||
| 64k | on the GPU | full logits | 192 | no | no | no | fits |
|
||||
|
||||
Reading:
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, with the CPU offload) fits 32k and 48k; the A100 and the RTX PRO 6000 do not. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is not known). The memory test shows it at once: if `seq_len 16000` already fails on an H200, the loss is the cause.
|
||||
Fallback (not built yet, small): compute the hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour) is only possible with the CPU offload and a chunked loss** (about 65 GB estimated at 48k: no margin). H200 141 GB (5 USD/hour) fits with margin;
|
||||
RTX PRO 6000 96 GB (2.75 USD/hour) fits with the offload and the chunked loss (est. 65 GB).
|
||||
- Samples over 32k are a minority (p90 39560 tokens in the current data, p50 24.8k): if 48k does not fit, 32k would drop 15 of 52 samples (29 %; 10 of them DDLS), so test 48k first.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): possible only with model parallelism (`device_map`), one card works at a time, Unsloth multi-GPU is limited; not recommended before the single card test.
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
|
||||
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about 65 GB at 48k, 67 GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
|
||||
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
|
||||
|
||||
## Proposed order
|
||||
1. Memory test on **h200** (about 30 minutes, 2.5 USD): sweep 16000,32000,48000 with the offload. If 48k fits with margin: stay on H200 or try rtx-pro-6000 for the real run.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours of work), repeat the test.
|
||||
3. Real run after the stage 1 : stage 2 ratio decision (below).
|
||||
## Order when the test is run
|
||||
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
|
||||
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
|
||||
|
||||
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 is faster, H200 about 2 to 2.5 times an A100)
|
||||
Tokens of the proposed run (1 epoch stage 1 and 3 epochs of the 47 stage 2 samples as built today: 2.53 M + 3 x 1.28 M = 6.4 M tokens): A100 about 5.9 h (about 15 USD), H200 about 2.7 h (about 14 USD). The real number comes from the memory test (`step_seconds`).
|
||||
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
|
||||
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
|
||||
A100 about 3.1 h (about 8 USD), H200 about 1.4 h (about 7 USD).
|
||||
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
|
||||
H200 about 4.3 h (about 21 USD), A100 about 9 h (about 23 USD). The real number comes from the memory test (`step_seconds`).
|
||||
|
||||
29
docs/foreign-objects-report.md
Normal file
29
docs/foreign-objects-report.md
Normal file
@@ -0,0 +1,29 @@
|
||||
# Other runs' objects in tool results: which runs are affected (2026-10-06)
|
||||
|
||||
Cause: the model sometimes names its own helper or test class `ZCL_<run prefix>_...` (prefix inside the name). The teardown looked only for names that start with the prefix, so such
|
||||
classes stayed in A4H and showed up in `sap_inactive_objects`, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, `harness/sweep.py`).
|
||||
Scan: `train/foreign_scan.py` (read only; result `runs/analysis/foreign_scan.json`). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object.
|
||||
|
||||
| group | runs | with foreign names in a tool result | with a foreign read | tools |
|
||||
|---|---|---|---|---|
|
||||
| official baseline Qwen (MacBook, 11 tasks) | 11 | 0 | 0 | - |
|
||||
| Devstral baseline (dropped) | 11 | 0 | 0 | - |
|
||||
| older Qwen baselines (archive) | 8 | 0 | 0 | - |
|
||||
| smoke | 1 | 0 | 0 | - |
|
||||
| eval empirical filter (DeepSeek on eval candidates) | 178 | 25 | 4 | sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5 |
|
||||
| eval generation validation (oracle, null, mutants) | 1344 | 0 | 0 | - |
|
||||
| early pilot runs | 16 | 1 | 0 | sap_inactive_objects 1 |
|
||||
| training trajectories (DeepSeek) | 117 | 67 | 15 | sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1 |
|
||||
| series A (local Qwen) | 20 | 3 | 0 | sap_inactive_objects 2, sap_search_object 1 |
|
||||
|
||||
## Reading
|
||||
- **Official baseline (Qwen, 11 tasks, 2026-10-04): not affected** (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers,
|
||||
and none of the 11 baseline runs called `sap_inactive_objects` (checked: 0 calls). **The baseline numbers stay as they are. Nothing was changed.**
|
||||
- **Eval generation (1344 validation runs: oracle, null, mutants): not affected** (they call no list or search tool).
|
||||
- **Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one**
|
||||
(G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there
|
||||
was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs).
|
||||
- **Early pilot runs:** 1 of 16 saw a foreign name in an inactive list.
|
||||
- **Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one.** Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool
|
||||
(`train/requeue_foreign.py`, rows in `runs/traj/summary_excluded.jsonl`) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder.
|
||||
- **Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads.** The scores (3 of 20) are not affected by a read; the series ran before the fix.
|
||||
@@ -63,6 +63,15 @@ Second attempts are part of the 130 runs. If the new reset gives 60 usage, this
|
||||
- Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
|
||||
- Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit.
|
||||
|
||||
## 5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks)
|
||||
1. `git status` clean and pushed; `python3 -m harness.restart_plan --panel <value>` prints the numbers (panel value from Kral, expected 0 after the reset).
|
||||
2. Services: A4H up, MCP answers, `python3 scripts_probe/lockprobe.py` shows no stale lock (only ATC runtime and debugger listener entries), `python3 -m harness.sweep` shows no leftover.
|
||||
3. `.env`: `BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD`, `BUDGET_RESERVE_USD=8`; no `STOP` flag, no `STOPPED.txt`; no running controller (`runs/pipeline/controller.lock`).
|
||||
4. Order of the first hour: step 0 eval slots for the new kinds (`evalset plan-new`, `run-new`, `run-new-k`), then the pending tasks of the kinds below target. The 9 trajectories that were
|
||||
moved back (`runs/traj/summary_excluded.jsonl`: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order.
|
||||
5. Settings to confirm: workers 2, `STREAM_GUARD` unset for the first runs and then 9000 for 5 DDLS/PROG runs (`docs/empty-response.md`), tool budget 100 for CDS.
|
||||
6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host.
|
||||
|
||||
## 6. After the restart
|
||||
|
||||
Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Stage 2 training set, build of 2026-10-06 07:36:30
|
||||
# Stage 2 training set, build of 2026-10-06 09:03:41
|
||||
|
||||
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
||||
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||
@@ -8,9 +8,7 @@ HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectorie
|
||||
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
|
||||
Dropped:
|
||||
- FUNC: {'read_of_another_runs_object': 3}
|
||||
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
|
||||
- TABL: {'read_of_another_runs_object': 1}
|
||||
- DDLS: {'over_48k': 5}
|
||||
- INTF: {'over_48k': 1}
|
||||
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
||||
|
||||
@@ -34,7 +32,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
|
||||
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
|
||||
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
@@ -52,7 +50,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
|
||||
|
||||
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
|
||||
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
|
||||
|
||||
| class of the trajectory | weight | reason |
|
||||
|---|---|---|
|
||||
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
|
||||
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
|
||||
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
|
||||
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
|
||||
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
|
||||
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
|
||||
|
||||
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
|
||||
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
|
||||
|
||||
## Other decisions of 2026-10-06
|
||||
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
|
||||
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
|
||||
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
|
||||
- **Memory test:** waits until the data is near the size of the first SFT run.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
|
||||
Reference in New Issue
Block a user