Compare commits

...

2 Commits

11 changed files with 422 additions and 41 deletions

View File

@@ -1,21 +1,24 @@
# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run) # bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
**No GPU job is started without Kral's go.** Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000` is the memory test, no data needed, 3 optimizer steps per length, stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
## Model (from config.json of Qwen3.8-27B) ## Model (from config.json of Qwen3.8-27B)
64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**. 64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**.
## Memory by component at 48k tokens (GB; the estimate, not a measurement) ## Memory by component (GB; estimate, not a measurement)
| component | GB | how | | component | 48k | 64k | how |
|---|---|---| |---|---|---|---|
| weights bf16 | 51.8 | 27.8 B x 2 bytes | | weights bf16 | 51.8 | 51.8 | 27.8 B x 2 bytes |
| LoRA weights + grads + 8-bit Adam | 1.0 | 107 M x 10 bytes | | LoRA weights + grads + 8-bit Adam | 1.0 | 1.0 | 107 M x 10 bytes |
| layer inputs for the backward pass (checkpoints) | 29.3 on the GPU, about 0 with the Unsloth CPU offload | 48k x 5120 x 2 bytes x 64 layers; the offload needs 29 GB of host RAM (all flavors have 142 GB or more) | | layer inputs for the backward pass | 29.3 on the GPU, about 0 with the Unsloth CPU offload (then 29 GB host RAM) | 39.1 / about 0 (host RAM 39 GB) | seq x 5120 x 2 bytes x 64 layers |
| recompute peak of one layer | 9.2 | MLP tensors 48k x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) | | recompute peak of one layer | 9.2 | 11.3 | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
| logits and loss | 3 chunked, 67 with full logits | full logits: 48k x 248320 x (bf16 + fp32 upcast + grad); only a chunked or fused cross entropy is realistic | | logits and loss | 3 chunked, 67 full logits | 3 chunked, 89 full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
## Fit by GPU (92 % of the card counted as usable) ## Fit by GPU (92 % of the card counted as usable)
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (needs model parallel) | | seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (model parallel) |
|---|---|---|---|---|---|---|---| |---|---|---|---|---|---|---|---|
| 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits | | 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits |
| 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits | | 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits |
@@ -29,20 +32,26 @@
| 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits | | 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits |
| 48k | on the GPU | chunked | 94 | no | tight | fits | fits | | 48k | on the GPU | chunked | 94 | no | tight | fits | fits |
| 48k | on the GPU | full logits | 158 | no | no | no | fits | | 48k | on the GPU | full logits | 158 | no | no | no | fits |
| 64k | CPU offload (Unsloth) | chunked | 67 | fits | fits | fits | fits |
| 64k | CPU offload (Unsloth) | full logits | 153 | no | no | no | fits |
| 64k | on the GPU | chunked | 106 | no | no | fits | fits |
| 64k | on the GPU | full logits | 192 | no | no | no | fits |
Reading: Reading:
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, with the CPU offload) fits 32k and 48k; the A100 and the RTX PRO 6000 do not. The run needs a fused or chunked cross entropy - **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is not known). The memory test shows it at once: if `seq_len 16000` already fails on an H200, the loss is the cause. (Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
Fallback (not built yet, small): compute the hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k. Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
- **A100 80 GB (2.50 USD/hour) is only possible with the CPU offload and a chunked loss** (about 65 GB estimated at 48k: no margin). H200 141 GB (5 USD/hour) fits with margin; - **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about 65 GB at 48k, 67 GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
RTX PRO 6000 96 GB (2.75 USD/hour) fits with the offload and the chunked loss (est. 65 GB). **RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
- Samples over 32k are a minority (p90 39560 tokens in the current data, p50 24.8k): if 48k does not fit, 32k would drop 15 of 52 samples (29 %; 10 of them DDLS), so test 48k first. - Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
- Multi-GPU (`a100x4`, `h200x2`): possible only with model parallelism (`device_map`), one card works at a time, Unsloth multi-GPU is limited; not recommended before the single card test.
## Proposed order ## Order when the test is run
1. Memory test on **h200** (about 30 minutes, 2.5 USD): sweep 16000,32000,48000 with the offload. If 48k fits with margin: stay on H200 or try rtx-pro-6000 for the real run. 1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
2. If the loss is the problem: build the chunked loss (about 2 hours of work), repeat the test. 2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
3. Real run after the stage 1 : stage 2 ratio decision (below). 3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 is faster, H200 about 2 to 2.5 times an A100) ## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
Tokens of the proposed run (1 epoch stage 1 and 3 epochs of the 47 stage 2 samples as built today: 2.53 M + 3 x 1.28 M = 6.4 M tokens): A100 about 5.9 h (about 15 USD), H200 about 2.7 h (about 14 USD). The real number comes from the memory test (`step_seconds`). Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
A100 about 3.1 h (about 8 USD), H200 about 1.4 h (about 7 USD).
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
H200 about 4.3 h (about 21 USD), A100 about 9 h (about 23 USD). The real number comes from the memory test (`step_seconds`).

View File

@@ -0,0 +1,29 @@
# Other runs' objects in tool results: which runs are affected (2026-10-06)
Cause: the model sometimes names its own helper or test class `ZCL_<run prefix>_...` (prefix inside the name). The teardown looked only for names that start with the prefix, so such
classes stayed in A4H and showed up in `sap_inactive_objects`, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, `harness/sweep.py`).
Scan: `train/foreign_scan.py` (read only; result `runs/analysis/foreign_scan.json`). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object.
| group | runs | with foreign names in a tool result | with a foreign read | tools |
|---|---|---|---|---|
| official baseline Qwen (MacBook, 11 tasks) | 11 | 0 | 0 | - |
| Devstral baseline (dropped) | 11 | 0 | 0 | - |
| older Qwen baselines (archive) | 8 | 0 | 0 | - |
| smoke | 1 | 0 | 0 | - |
| eval empirical filter (DeepSeek on eval candidates) | 178 | 25 | 4 | sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5 |
| eval generation validation (oracle, null, mutants) | 1344 | 0 | 0 | - |
| early pilot runs | 16 | 1 | 0 | sap_inactive_objects 1 |
| training trajectories (DeepSeek) | 117 | 67 | 15 | sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1 |
| series A (local Qwen) | 20 | 3 | 0 | sap_inactive_objects 2, sap_search_object 1 |
## Reading
- **Official baseline (Qwen, 11 tasks, 2026-10-04): not affected** (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers,
and none of the 11 baseline runs called `sap_inactive_objects` (checked: 0 calls). **The baseline numbers stay as they are. Nothing was changed.**
- **Eval generation (1344 validation runs: oracle, null, mutants): not affected** (they call no list or search tool).
- **Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one**
(G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there
was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs).
- **Early pilot runs:** 1 of 16 saw a foreign name in an inactive list.
- **Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one.** Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool
(`train/requeue_foreign.py`, rows in `runs/traj/summary_excluded.jsonl`) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder.
- **Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads.** The scores (3 of 20) are not affected by a read; the series ran before the fix.

20
docs/own-test-mutation.md Normal file
View File

@@ -0,0 +1,20 @@
# Own-test mutation score of the accepted trajectories (item D, metadata only)
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
## Result
90 accepted trajectories looked at: {'scored': 63, 'not_supported': 5, 'no_own_tests': 22}.
- Scored: 63; the model's tests also passed on the correct reference: 53 (the others are not reliable: a test that fails on the correct solution kills every mutant).
- Mean score over the reliable ones: 0.99; median 1.00; distribution {'1.0': 50, '>=0.75': 3} (n = 53).
- By kind (mean, n): {CLAS: 1.00 (45), DDLS: 0.90 (5), FUNC: 1.00 (3)}
- `no_own_tests`: 22 trajectories were accepted without any own test (at most 85 points); `no_mutants`: 0; `not_supported` (PROG: tests are inside the program): 5.
## Weakest (score under 0.75, tests pass on the reference)
## Use
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).

View File

@@ -63,6 +63,15 @@ Second attempts are part of the 130 runs. If the new reset gives 60 usage, this
- Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k. - Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
- Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit. - Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit.
## 5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks)
1. `git status` clean and pushed; `python3 -m harness.restart_plan --panel <value>` prints the numbers (panel value from Kral, expected 0 after the reset).
2. Services: A4H up, MCP answers, `python3 scripts_probe/lockprobe.py` shows no stale lock (only ATC runtime and debugger listener entries), `python3 -m harness.sweep` shows no leftover.
3. `.env`: `BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD`, `BUDGET_RESERVE_USD=8`; no `STOP` flag, no `STOPPED.txt`; no running controller (`runs/pipeline/controller.lock`).
4. Order of the first hour: step 0 eval slots for the new kinds (`evalset plan-new`, `run-new`, `run-new-k`), then the pending tasks of the kinds below target. The 9 trajectories that were
moved back (`runs/traj/summary_excluded.jsonl`: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order.
5. Settings to confirm: workers 2, `STREAM_GUARD` unset for the first runs and then 9000 for 5 DDLS/PROG runs (`docs/empty-response.md`), tool budget 100 for CDS.
6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host.
## 6. After the restart ## 6. After the restart
Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the

View File

@@ -1,4 +1,4 @@
# Stage 2 training set, build of 2026-10-06 06:15:28 # Stage 2 training set, build of 2026-10-06 09:03:41
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`. Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read. HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
@@ -8,9 +8,7 @@ HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectorie
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`) -> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86). -> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
Dropped: Dropped:
- FUNC: {'read_of_another_runs_object': 3} - DDLS: {'over_48k': 5}
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
- TABL: {'read_of_another_runs_object': 1}
- INTF: {'over_48k': 1} - INTF: {'over_48k': 1}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary. Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
@@ -34,7 +32,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt. - Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60). - Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides) ## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data): Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* | | stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
@@ -52,7 +50,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again. regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints. (3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions. (4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1. **Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
| class of the trajectory | weight | reason |
|---|---|---|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
## Other decisions of 2026-10-06
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
- **Memory test:** waits until the data is near the size of the first SFT run.
## Also built ## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights). `train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).

View File

@@ -43,7 +43,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt. - Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60). - Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides) ## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data): Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* | | stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
@@ -61,7 +61,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again. regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints. (3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions. (4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1. **Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
| class of the trajectory | weight | reason |
|---|---|---|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
## Other decisions of 2026-10-06
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
- **Memory test:** waits until the data is near the size of the first SFT run.
## Also built ## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights). `train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).

79
train/foreign_scan.py Normal file
View File

@@ -0,0 +1,79 @@
"""Read-only scan: which runs saw other runs' objects in tool results, or read them (teardown/proxy bug of 2026-10-06).
python3 train/foreign_scan.py -> prints a summary and writes runs/analysis/foreign_scan.json
"""
import glob
import json
import os
import re
import sys
from collections import Counter
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
from harness.task import prefix_for # noqa: E402
START = re.compile(r"Z\d[0-9A-Z]{6}_")
MID_NAME = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
GROUPS = [("official baseline Qwen (MacBook, 11 tasks)", "runs/stage1/baseline_qwen/*"),
("Devstral baseline (dropped)", "runs/archive/devstral/baseline_devstral/*"),
("older Qwen baselines (archive)", "runs/archive/qwen38/*/*"),
("smoke", "runs/stage1/smoke/*"),
("eval empirical filter (DeepSeek on eval candidates)", "runs/emp/*"),
("eval generation validation (oracle, null, mutants)", "runs/gen/*"),
("early pilot runs", "runs/2*_*"),
("training trajectories (DeepSeek)", "runs/traj/*_G*"),
("series A (local Qwen)", "runs/local_qwen/runs/*")]
def own_prefix(d):
m = re.match(r"^(\d+)_([GT]\d+)_", os.path.basename(d))
return prefix_for(int(m.group(1)), m.group(2)).upper() if m else None
def scan(d):
own = own_prefix(d)
p = os.path.join(d, "trajectory.jsonl")
if not own or not os.path.exists(p):
return None
seen, reads, tools = set(), 0, Counter()
for l in open(p):
try:
e = json.loads(l)
except ValueError:
continue
if "tool" not in e:
continue
text = (e.get("result") or "").upper()
found = {x for x in START.findall(text) if x != own and x[1].isdigit()}
found |= {m.group(1).upper() for m in re.finditer(r"\b[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", text) if m.group(1).upper() != own}
if found:
seen |= found
tools[e["tool"]] += 1
name = str((e.get("args") or {}).get("objectName", "")).upper()
m = START.match(name) or MID_NAME.match(name)
pre = (m.group(1) if m and m.re is MID_NAME else (m.group(0) if m else "")).upper()
if e["tool"] in READ and pre and pre != own:
reads += 1
return {"run": os.path.basename(d), "foreign_prefixes_seen": len(seen), "foreign_reads": reads, "tools": dict(tools)}
def main():
out = {}
for label, pat in GROUPS:
dirs = [d for d in sorted(glob.glob(os.path.join(ROOT, pat))) if os.path.isdir(d)]
rows = [r for r in (scan(d) for d in dirs) if r]
seen = [r for r in rows if r["foreign_prefixes_seen"]]
reads = [r for r in rows if r["foreign_reads"]]
out[label] = {"runs": len(rows), "runs_with_foreign_names": len(seen), "runs_with_foreign_reads": len(reads),
"tools": dict(sum((Counter(r["tools"]) for r in seen), Counter())),
"affected": [r["run"] for r in seen][:40], "read_runs": [r["run"] for r in reads]}
print(f"{label}: {len(rows)} runs | foreign names in tool results: {len(seen)} | foreign reads: {len(reads)} | tools {out[label]['tools']}")
os.makedirs(os.path.join(ROOT, "runs", "analysis"), exist_ok=True)
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "foreign_scan.json"), "w"), indent=1)
if __name__ == "__main__":
main()

View File

@@ -5,10 +5,10 @@
"""Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go. """Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go.
Memory test first (a few dollars, finds the largest sequence length that fits, no data needed): Memory test first (a few dollars, finds the largest sequence length that fits, no data needed):
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000 hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000,64000
Real run (after the memory test and Kral's decision on the ratio): Real run (after the memory test and Kral's decision on the ratio):
hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\ hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s1-epochs 1 --s2-epochs 3 --out erhankeseli/abap-mixed-adapter --stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s2-epochs 3 --s2-loss-share 0.6 --out erhankeseli/abap-mixed-adapter
Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns; Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns;
none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut. none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut.
@@ -35,6 +35,18 @@ def tokenize_masked(tok, text, spans, max_len):
return ids, labels return ids, labels
def s1_epochs_for(share, s2_loss_tokens, s2_epochs, s1_tokens):
"""Epochs of stage 1 so that stage 2 carries `share` of all loss tokens (Kral + Opus 2026-10-06: share = 0.6)."""
s2_total = s2_loss_tokens * s2_epochs
return (1 - share) / share * s2_total / max(s1_tokens, 1)
def copies(weight, rnd):
"""Weight 1 = one copy per epoch, 2 = two, 0.5 = a copy in half of the epochs (a down-weight, not a drop)."""
w = max(float(weight), 0.0)
return int(w) + (1 if rnd.random() < w - int(w) else 0)
def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed): def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
rnd = random.Random(seed) rnd = random.Random(seed)
rows = [] rows = []
@@ -45,7 +57,7 @@ def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))] rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))]
for ep in range(int(s2_epochs)): for ep in range(int(s2_epochs)):
for r in stage2: for r in stage2:
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * max(1, round(float(r.get("weight", 1.0)))) rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * copies(r.get("weight", 1.0), rnd)
rnd.shuffle(rows) rnd.shuffle(rows)
out, skipped = [], {"s1": 0, "s2": 0} out, skipped = [], {"s1": 0, "s2": 0}
for src, text, spans, _ in rows: for src, text, spans, _ in rows:
@@ -65,8 +77,9 @@ def main():
ap.add_argument("--s1-file", default="train.jsonl") ap.add_argument("--s1-file", default="train.jsonl")
ap.add_argument("--s2-file", default="stage2_train.jsonl") ap.add_argument("--s2-file", default="stage2_train.jsonl")
ap.add_argument("--s2-valid", default="stage2_valid.jsonl") ap.add_argument("--s2-valid", default="stage2_valid.jsonl")
ap.add_argument("--s1-epochs", type=float, default=1.0) ap.add_argument("--s1-epochs", type=float, default=None, help="epochs of stage 1; default: computed from --s2-loss-share")
ap.add_argument("--s2-epochs", type=float, default=3.0) ap.add_argument("--s2-epochs", type=float, default=3.0)
ap.add_argument("--s2-loss-share", type=float, default=0.6, help="share of the loss tokens that stage 2 carries (decision 2026-10-06: 0.6)")
ap.add_argument("--max-seq", type=int, default=48000) ap.add_argument("--max-seq", type=int, default=48000)
ap.add_argument("--rank", type=int, default=16) ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--alpha", type=int, default=32) ap.add_argument("--alpha", type=int, default=32)
@@ -77,7 +90,7 @@ def main():
ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter") ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter")
ap.add_argument("--save-every", type=int, default=0) ap.add_argument("--save-every", type=int, default=0)
ap.add_argument("--memory-test", action="store_true") ap.add_argument("--memory-test", action="store_true")
ap.add_argument("--sweep", default="16000,32000,48000", help="memory test: sequence lengths, tried in this order, stops at the first OOM") ap.add_argument("--sweep", default="16000,32000,48000,64000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length") ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length")
a = ap.parse_args() a = ap.parse_args()
@@ -85,7 +98,8 @@ def main():
from unsloth import FastLanguageModel from unsloth import FastLanguageModel
from huggingface_hub import HfApi, hf_hub_download from huggingface_hub import HfApi, hf_hub_download
token = os.environ["HF_TOKEN"] token = os.environ["HF_TOKEN"]
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=a.max_seq, load_in_4bit=False, dtype=torch.bfloat16, token=token) seq_for_model = max([a.max_seq] + ([int(x) for x in a.sweep.split(",")] if a.memory_test else []))
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=seq_for_model, load_in_4bit=False, dtype=torch.bfloat16, token=token)
model = FastLanguageModel.get_peft_model( model = FastLanguageModel.get_peft_model(
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none", model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"], target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"],
@@ -132,6 +146,11 @@ def main():
p = hf_hub_download(repo, fname, repo_type="dataset", token=token) p = hf_hub_download(repo, fname, repo_type="dataset", token=token)
return [json.loads(l) for l in open(p)] return [json.loads(l) for l in open(p)]
s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file) s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file)
if a.s1_epochs is None:
s1_tok = sum(len(tok(r["text"], add_special_tokens=False)["input_ids"]) for r in s1)
s2_loss = sum(r["assistant_tokens"] * float(r.get("weight", 1.0)) for r in s2)
a.s1_epochs = s1_epochs_for(a.s2_loss_share, s2_loss, a.s2_epochs, s1_tok)
print("STAGE1 EPOCHS", round(a.s1_epochs, 3), "(stage 1 tokens", s1_tok, ", stage 2 loss tokens per epoch", round(s2_loss), ", share", a.s2_loss_share, ")", flush=True)
ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006) ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006)
loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex) loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex)
print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens", print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens",

View File

@@ -10,17 +10,47 @@ def identity(row):
return 1.0 return 1.0
# Weights from the own-test mutation score (Kral + Opus 2026-10-06: down-weight, do not drop). Weight w means: w copies per epoch, a fraction is a
# copy in that share of the epochs (0.5 = every second epoch on average). The acceptance filter itself is unchanged.
OWN_TEST_WEIGHTS = {
"reliable, score >= 0.75": 1.0, # the own tests pass on the correct reference and kill at least 3 of 4 mutants
"reliable, score 0.5 to 0.75": 0.75,
"reliable, score < 0.5": 0.5,
"unreliable (own tests fail on the correct reference)": 0.5, # they may encode model specific behavior
"no own tests": 0.5, # accepted at 85 points at most; the habit of writing tests is part of the behavior we want
"no signal (PROG, no mutants, not scored)": 1.0,
}
def own_test_class(m):
if m is None:
return "no signal (PROG, no mutants, not scored)"
st = m.get("status")
if st == "no_own_tests":
return "no own tests"
if st != "scored":
return "no signal (PROG, no mutants, not scored)"
if not m.get("tests_pass_on_reference"):
return "unreliable (own tests fail on the correct reference)"
s = m.get("score")
if s is None:
return "no signal (PROG, no mutants, not scored)"
return "reliable, score >= 0.75" if s >= 0.75 else "reliable, score 0.5 to 0.75" if s >= 0.5 else "reliable, score < 0.5"
def own_test_weight(row): def own_test_weight(row):
"""Item D (own-test mutation score, metadata only for now): reads runs/traj/<run>/own_test_mutation.json when it exists, """Item D: reads runs/traj/<run>/own_test_mutation.json, stores the score as extra data and sets the weight by OWN_TEST_WEIGHTS."""
stores it as extra data and does NOT drop or reweight (Kral + Opus 2026-10-06: do not change the acceptance yet)."""
run = row["id"].split("_r")[-1] run = row["id"].split("_r")[-1]
m = None
for d in os.listdir(os.path.join(ROOT, "runs", "traj")): for d in os.listdir(os.path.join(ROOT, "runs", "traj")):
if d.startswith(run + "_"): if d.startswith(run + "_"):
p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json") p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json")
if os.path.exists(p): if os.path.exists(p):
m = json.load(open(p)) m = json.load(open(p))
return {"keep": True, "weight": 1.0, "extra": {"own_test_mutation": m.get("score"), "own_test_mutants": m.get("mutants")}} break
return {"keep": True, "weight": 1.0} cls = own_test_class(m)
return {"keep": True, "weight": OWN_TEST_WEIGHTS[cls],
"extra": {"own_test_class": cls, "own_test_mutation": (m or {}).get("score"), "own_test_reliable": bool((m or {}).get("tests_pass_on_reference"))}}
def repair_up(row): def repair_up(row):

71
train/mem_table.py Normal file
View File

@@ -0,0 +1,71 @@
"""Writes docs/bf16-memory.md (estimates for the bf16 LoRA run of Qwen 3.8 27B by GPU and sequence length)."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
H, L, V, I = 5120, 64, 248320, 17408
r = 16
lora = 16 * r * ((H + 6144) + (H + 1024) + (H + 1024) + (6144 + H)) + 48 * r * ((H + 10240) + (H + 6144) + (6144 + H)) + L * r * ((H + I) * 2 + (I + H))
W = 27.8e9 * 2 / 2**30
lora_gb = lora * 10 / 2**30
def est(seq, offload, chunked):
ckpt = 0.0 if offload else seq * H * 2 * L / 2**30
layer = seq * I * 2 * 4 / 2**30 + 3
logits = (seq * V * 2 * 3 / 2**30) if not chunked else 3.0
return W + lora_gb + ckpt + layer + logits, dict(weights=W, lora=lora_gb, ckpt=ckpt, layer=layer, logits=logits)
gpus = [("a100-large", "1x A100 80 GB", 80), ("rtx-pro-6000", "1x RTX PRO 6000 96 GB", 96), ("h200", "1x H200 141 GB", 141), ("h200x2", "2x H200 282 GB (model parallel)", 282)]
tab = "| seq | checkpoints | loss | estimated peak GB | " + " | ".join(g[1] for g in gpus) + " |\n|---|---|---|---|" + "---|" * len(gpus) + "\n"
for seq in (16000, 32000, 48000, 64000):
for off in (True, False):
for ch in (True, False):
t, _ = est(seq, off, ch)
cells = ["fits" if t <= g[2] * 0.92 else ("tight" if t <= g[2] else "no") for g in gpus]
tab += f"| {seq // 1000}k | {'CPU offload (Unsloth)' if off else 'on the GPU'} | {'chunked' if ch else 'full logits'} | {t:.0f} | " + " | ".join(cells) + " |\n"
_, b48 = est(48000, True, True)
_, b64 = est(64000, True, True)
doc = f"""# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
## Model (from config.json of Qwen3.8-27B)
64 layers (48 linear attention, 16 full attention), hidden {H}, MLP {I}, vocab {V}, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **{lora / 1e6:.0f} M trainable parameters**.
## Memory by component (GB; estimate, not a measurement)
| component | 48k | 64k | how |
|---|---|---|---|
| weights bf16 | {b48['weights']:.1f} | {b64['weights']:.1f} | 27.8 B x 2 bytes |
| LoRA weights + grads + 8-bit Adam | {b48['lora']:.1f} | {b64['lora']:.1f} | {lora / 1e6:.0f} M x 10 bytes |
| layer inputs for the backward pass | {48000 * H * 2 * L / 2**30:.1f} on the GPU, about 0 with the Unsloth CPU offload (then {48000 * H * 2 * L / 2**30:.0f} GB host RAM) | {64000 * H * 2 * L / 2**30:.1f} / about 0 (host RAM {64000 * H * 2 * L / 2**30:.0f} GB) | seq x 5120 x 2 bytes x 64 layers |
| recompute peak of one layer | {b48['layer']:.1f} | {b64['layer']:.1f} | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
| logits and loss | {b48['logits']:.0f} chunked, {48000 * V * 2 * 3 / 2**30:.0f} full logits | {b64['logits']:.0f} chunked, {64000 * V * 2 * 3 / 2**30:.0f} full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
## Fit by GPU (92 % of the card counted as usable)
{tab}
Reading:
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about {b48['weights'] + b48['lora'] + b48['layer'] + b48['logits']:.0f} GB at 48k, {b64['weights'] + b64['lora'] + b64['layer'] + b64['logits']:.0f} GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
## Order when the test is run
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
## Time and cost (estimate from the nf4 run: {6750 / 30.5:.0f} tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
A100 about {3.3e6 / 300 / 3600:.1f} h (about {3.3e6 / 300 / 3600 * 2.5:.0f} USD), H200 about {3.3e6 / 650 / 3600:.1f} h (about {3.3e6 / 650 / 3600 * 5:.0f} USD).
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
H200 about {10e6 / 650 / 3600:.1f} h (about {10e6 / 650 / 3600 * 5:.0f} USD), A100 about {10e6 / 300 / 3600:.0f} h (about {10e6 / 300 / 3600 * 2.5:.0f} USD). The real number comes from the memory test (`step_seconds`).
"""
open(os.path.join(ROOT, "docs", "bf16-memory.md"), "w").write(doc)
print("written")

75
train/requeue_foreign.py Normal file
View File

@@ -0,0 +1,75 @@
"""Trajectories in which the model read another run's leftover object go back to the pending pool (Kral + Opus 2026-10-06).
python3 train/requeue_foreign.py dry run
python3 train/requeue_foreign.py --apply moves their rows from runs/traj/summary.jsonl to runs/traj/summary_excluded.jsonl
(backup: summary.jsonl.bak-<time>); the task then counts as not run and is run again
after the reset in the normal order (kind deficit). The run folders stay.
"""
import json
import os
import shutil
import sys
import time
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
from harness import mix # noqa: E402
START = __import__("re").compile(r"^(Z\d[0-9A-Z]{6}_)", __import__("re").I)
MID = __import__("re").compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", __import__("re").I)
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
def foreign_reads(rec):
own = rec["prefix"].upper()
n = 0
for m in rec["messages"]:
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
if c["function"]["name"] not in READ:
continue
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
name = str(a.get("objectName", "")).upper()
mm = START.match(name) or MID.match(name)
if mm and mm.group(1).upper() != own:
n += 1
return n
def main():
apply = "--apply" in sys.argv
path = os.path.join(ROOT, "runs", "traj", "summary.jsonl")
rows = [json.loads(l) for l in open(path)]
keep, out = [], []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
if os.path.exists(p):
rec = json.load(open(p))
if acc.judge(rec, r, 80)[0]:
n = foreign_reads(rec)
if n:
out.append(dict(r, excluded_reason="read of another run's object", foreign_reads=n, kind=mix.kind_of_task_dir(r["task"]),
excluded_at=time.strftime("%F %T")))
continue
keep.append(r)
print(len(out), "accepted trajectories with a foreign read:", [(o["task"], o["attempt"], o["kind"], o["foreign_reads"]) for o in out])
if not apply:
return
shutil.copy(path, path + ".bak-" + time.strftime("%Y%m%d-%H%M%S"))
with open(os.path.join(ROOT, "runs", "traj", "summary_excluded.jsonl"), "a") as f:
for o in out:
f.write(json.dumps(o) + "\n")
open(path, "w").write("".join(json.dumps(r) + "\n" for r in keep))
print("moved; summary.jsonl now", len(keep), "rows")
if __name__ == "__main__":
main()