diff --git a/docs/bf16-memory.md b/docs/bf16-memory.md index adf321f..4ea7cbe 100644 --- a/docs/bf16-memory.md +++ b/docs/bf16-memory.md @@ -1,21 +1,24 @@ # bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run) -**No GPU job is started without Kral's go.** Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000` is the memory test, no data needed, 3 optimizer steps per length, stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). +**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the +count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length, +stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000` +until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k). ## Model (from config.json of Qwen3.8-27B) 64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**. -## Memory by component at 48k tokens (GB; the estimate, not a measurement) -| component | GB | how | -|---|---|---| -| weights bf16 | 51.8 | 27.8 B x 2 bytes | -| LoRA weights + grads + 8-bit Adam | 1.0 | 107 M x 10 bytes | -| layer inputs for the backward pass (checkpoints) | 29.3 on the GPU, about 0 with the Unsloth CPU offload | 48k x 5120 x 2 bytes x 64 layers; the offload needs 29 GB of host RAM (all flavors have 142 GB or more) | -| recompute peak of one layer | 9.2 | MLP tensors 48k x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) | -| logits and loss | 3 chunked, 67 with full logits | full logits: 48k x 248320 x (bf16 + fp32 upcast + grad); only a chunked or fused cross entropy is realistic | +## Memory by component (GB; estimate, not a measurement) +| component | 48k | 64k | how | +|---|---|---|---| +| weights bf16 | 51.8 | 51.8 | 27.8 B x 2 bytes | +| LoRA weights + grads + 8-bit Adam | 1.0 | 1.0 | 107 M x 10 bytes | +| layer inputs for the backward pass | 29.3 on the GPU, about 0 with the Unsloth CPU offload (then 29 GB host RAM) | 39.1 / about 0 (host RAM 39 GB) | seq x 5120 x 2 bytes x 64 layers | +| recompute peak of one layer | 9.2 | 11.3 | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) | +| logits and loss | 3 chunked, 67 full logits | 3 chunked, 89 full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) | ## Fit by GPU (92 % of the card counted as usable) -| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (needs model parallel) | +| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (model parallel) | |---|---|---|---|---|---|---|---| | 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits | | 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits | @@ -29,20 +32,26 @@ | 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits | | 48k | on the GPU | chunked | 94 | no | tight | fits | fits | | 48k | on the GPU | full logits | 158 | no | no | no | fits | +| 64k | CPU offload (Unsloth) | chunked | 67 | fits | fits | fits | fits | +| 64k | CPU offload (Unsloth) | full logits | 153 | no | no | no | fits | +| 64k | on the GPU | chunked | 106 | no | no | fits | fits | +| 64k | on the GPU | full logits | 192 | no | no | no | fits | Reading: -- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, with the CPU offload) fits 32k and 48k; the A100 and the RTX PRO 6000 do not. The run needs a fused or chunked cross entropy - (Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is not known). The memory test shows it at once: if `seq_len 16000` already fails on an H200, the loss is the cause. - Fallback (not built yet, small): compute the hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k. -- **A100 80 GB (2.50 USD/hour) is only possible with the CPU offload and a chunked loss** (about 65 GB estimated at 48k: no margin). H200 141 GB (5 USD/hour) fits with margin; - RTX PRO 6000 96 GB (2.75 USD/hour) fits with the offload and the chunked loss (est. 65 GB). -- Samples over 32k are a minority (p90 39560 tokens in the current data, p50 24.8k): if 48k does not fit, 32k would drop 15 of 52 samples (29 %; 10 of them DDLS), so test 48k first. -- Multi-GPU (`a100x4`, `h200x2`): possible only with model parallelism (`device_map`), one card works at a time, Unsloth multi-GPU is limited; not recommended before the single card test. +- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy + (Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause. + Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k. +- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about 65 GB at 48k, 67 GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin; + **RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss. +- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test. -## Proposed order -1. Memory test on **h200** (about 30 minutes, 2.5 USD): sweep 16000,32000,48000 with the offload. If 48k fits with margin: stay on H200 or try rtx-pro-6000 for the real run. -2. If the loss is the problem: build the chunked loss (about 2 hours of work), repeat the test. -3. Real run after the stage 1 : stage 2 ratio decision (below). +## Order when the test is run +1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin. +2. If the loss is the problem: build the chunked loss (about 2 hours), repeat. +3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied). -## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 is faster, H200 about 2 to 2.5 times an A100) -Tokens of the proposed run (1 epoch stage 1 and 3 epochs of the 47 stage 2 samples as built today: 2.53 M + 3 x 1.28 M = 6.4 M tokens): A100 about 5.9 h (about 15 USD), H200 about 2.7 h (about 14 USD). The real number comes from the memory test (`step_seconds`). +## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100) +Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total. +A100 about 3.1 h (about 8 USD), H200 about 1.4 h (about 7 USD). +With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens, +H200 about 4.3 h (about 21 USD), A100 about 9 h (about 23 USD). The real number comes from the memory test (`step_seconds`). diff --git a/docs/foreign-objects-report.md b/docs/foreign-objects-report.md new file mode 100644 index 0000000..3aa5fce --- /dev/null +++ b/docs/foreign-objects-report.md @@ -0,0 +1,29 @@ +# Other runs' objects in tool results: which runs are affected (2026-10-06) + +Cause: the model sometimes names its own helper or test class `ZCL__...` (prefix inside the name). The teardown looked only for names that start with the prefix, so such +classes stayed in A4H and showed up in `sap_inactive_objects`, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, `harness/sweep.py`). +Scan: `train/foreign_scan.py` (read only; result `runs/analysis/foreign_scan.json`). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object. + +| group | runs | with foreign names in a tool result | with a foreign read | tools | +|---|---|---|---|---| +| official baseline Qwen (MacBook, 11 tasks) | 11 | 0 | 0 | - | +| Devstral baseline (dropped) | 11 | 0 | 0 | - | +| older Qwen baselines (archive) | 8 | 0 | 0 | - | +| smoke | 1 | 0 | 0 | - | +| eval empirical filter (DeepSeek on eval candidates) | 178 | 25 | 4 | sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5 | +| eval generation validation (oracle, null, mutants) | 1344 | 0 | 0 | - | +| early pilot runs | 16 | 1 | 0 | sap_inactive_objects 1 | +| training trajectories (DeepSeek) | 117 | 67 | 15 | sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1 | +| series A (local Qwen) | 20 | 3 | 0 | sap_inactive_objects 2, sap_search_object 1 | + +## Reading +- **Official baseline (Qwen, 11 tasks, 2026-10-04): not affected** (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers, + and none of the 11 baseline runs called `sap_inactive_objects` (checked: 0 calls). **The baseline numbers stay as they are. Nothing was changed.** +- **Eval generation (1344 validation runs: oracle, null, mutants): not affected** (they call no list or search tool). +- **Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one** + (G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there + was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs). +- **Early pilot runs:** 1 of 16 saw a foreign name in an inactive list. +- **Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one.** Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool + (`train/requeue_foreign.py`, rows in `runs/traj/summary_excluded.jsonl`) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder. +- **Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads.** The scores (3 of 20) are not affected by a read; the series ran before the fix. diff --git a/docs/restart-12-oktober.md b/docs/restart-12-oktober.md index 109d3b0..e137d6c 100644 --- a/docs/restart-12-oktober.md +++ b/docs/restart-12-oktober.md @@ -63,6 +63,15 @@ Second attempts are part of the 130 runs. If the new reset gives 60 usage, this - Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k. - Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit. +## 5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks) +1. `git status` clean and pushed; `python3 -m harness.restart_plan --panel ` prints the numbers (panel value from Kral, expected 0 after the reset). +2. Services: A4H up, MCP answers, `python3 scripts_probe/lockprobe.py` shows no stale lock (only ATC runtime and debugger listener entries), `python3 -m harness.sweep` shows no leftover. +3. `.env`: `BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD`, `BUDGET_RESERVE_USD=8`; no `STOP` flag, no `STOPPED.txt`; no running controller (`runs/pipeline/controller.lock`). +4. Order of the first hour: step 0 eval slots for the new kinds (`evalset plan-new`, `run-new`, `run-new-k`), then the pending tasks of the kinds below target. The 9 trajectories that were + moved back (`runs/traj/summary_excluded.jsonl`: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order. +5. Settings to confirm: workers 2, `STREAM_GUARD` unset for the first runs and then 9000 for 5 DDLS/PROG runs (`docs/empty-response.md`), tool budget 100 for CDS. +6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host. + ## 6. After the restart Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the diff --git a/docs/stage2-build.md b/docs/stage2-build.md index 2310541..3ac6f19 100644 --- a/docs/stage2-build.md +++ b/docs/stage2-build.md @@ -1,4 +1,4 @@ -# Stage 2 training set, build of 2026-10-06 07:36:30 +# Stage 2 training set, build of 2026-10-06 09:03:41 Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`. HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read. @@ -8,9 +8,7 @@ HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectorie -> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`) -> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86). Dropped: -- FUNC: {'read_of_another_runs_object': 3} -- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5} -- TABL: {'read_of_another_runs_object': 1} +- DDLS: {'over_48k': 5} - INTF: {'over_48k': 1} Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary. @@ -34,7 +32,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys - Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt. - Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60). -## Stage 1 : stage 2 ratio (proposal, Kral decides) +## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule) Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data): | stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* | @@ -52,7 +50,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again. (3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints. (4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions. -Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1. +**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens. + +## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`) +A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged. + +| class of the trajectory | weight | reason | +|---|---|---| +| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want | +| reliable, score 0.5 to 0.75 | 0.75 | weaker tests | +| reliable, score under 0.5 | 0.5 | tests that miss most faults | +| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior | +| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach | +| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad | + +In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1. +Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look. + +## Other decisions of 2026-10-06 +- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve). +- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then. +- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`). +- **Memory test:** waits until the data is near the size of the first SFT run. ## Also built `train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights). diff --git a/train/build_doc.py b/train/build_doc.py index f7bf284..52e2f11 100644 --- a/train/build_doc.py +++ b/train/build_doc.py @@ -43,7 +43,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys - Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt. - Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60). -## Stage 1 : stage 2 ratio (proposal, Kral decides) +## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule) Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data): | stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* | @@ -61,7 +61,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again. (3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints. (4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions. -Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1. +**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens. + +## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`) +A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged. + +| class of the trajectory | weight | reason | +|---|---|---| +| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want | +| reliable, score 0.5 to 0.75 | 0.75 | weaker tests | +| reliable, score under 0.5 | 0.5 | tests that miss most faults | +| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior | +| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach | +| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad | + +In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1. +Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look. + +## Other decisions of 2026-10-06 +- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve). +- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then. +- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`). +- **Memory test:** waits until the data is near the size of the first SFT run. ## Also built `train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights). diff --git a/train/foreign_scan.py b/train/foreign_scan.py new file mode 100644 index 0000000..f3f5e1b --- /dev/null +++ b/train/foreign_scan.py @@ -0,0 +1,79 @@ +"""Read-only scan: which runs saw other runs' objects in tool results, or read them (teardown/proxy bug of 2026-10-06). + + python3 train/foreign_scan.py -> prints a summary and writes runs/analysis/foreign_scan.json +""" +import glob +import json +import os +import re +import sys +from collections import Counter + +ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +sys.path.insert(0, ROOT) +from harness.task import prefix_for # noqa: E402 + +START = re.compile(r"Z\d[0-9A-Z]{6}_") +MID_NAME = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I) +READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object", + "sap_syntax_check", "sap_atc_run") +GROUPS = [("official baseline Qwen (MacBook, 11 tasks)", "runs/stage1/baseline_qwen/*"), + ("Devstral baseline (dropped)", "runs/archive/devstral/baseline_devstral/*"), + ("older Qwen baselines (archive)", "runs/archive/qwen38/*/*"), + ("smoke", "runs/stage1/smoke/*"), + ("eval empirical filter (DeepSeek on eval candidates)", "runs/emp/*"), + ("eval generation validation (oracle, null, mutants)", "runs/gen/*"), + ("early pilot runs", "runs/2*_*"), + ("training trajectories (DeepSeek)", "runs/traj/*_G*"), + ("series A (local Qwen)", "runs/local_qwen/runs/*")] + + +def own_prefix(d): + m = re.match(r"^(\d+)_([GT]\d+)_", os.path.basename(d)) + return prefix_for(int(m.group(1)), m.group(2)).upper() if m else None + + +def scan(d): + own = own_prefix(d) + p = os.path.join(d, "trajectory.jsonl") + if not own or not os.path.exists(p): + return None + seen, reads, tools = set(), 0, Counter() + for l in open(p): + try: + e = json.loads(l) + except ValueError: + continue + if "tool" not in e: + continue + text = (e.get("result") or "").upper() + found = {x for x in START.findall(text) if x != own and x[1].isdigit()} + found |= {m.group(1).upper() for m in re.finditer(r"\b[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", text) if m.group(1).upper() != own} + if found: + seen |= found + tools[e["tool"]] += 1 + name = str((e.get("args") or {}).get("objectName", "")).upper() + m = START.match(name) or MID_NAME.match(name) + pre = (m.group(1) if m and m.re is MID_NAME else (m.group(0) if m else "")).upper() + if e["tool"] in READ and pre and pre != own: + reads += 1 + return {"run": os.path.basename(d), "foreign_prefixes_seen": len(seen), "foreign_reads": reads, "tools": dict(tools)} + + +def main(): + out = {} + for label, pat in GROUPS: + dirs = [d for d in sorted(glob.glob(os.path.join(ROOT, pat))) if os.path.isdir(d)] + rows = [r for r in (scan(d) for d in dirs) if r] + seen = [r for r in rows if r["foreign_prefixes_seen"]] + reads = [r for r in rows if r["foreign_reads"]] + out[label] = {"runs": len(rows), "runs_with_foreign_names": len(seen), "runs_with_foreign_reads": len(reads), + "tools": dict(sum((Counter(r["tools"]) for r in seen), Counter())), + "affected": [r["run"] for r in seen][:40], "read_runs": [r["run"] for r in reads]} + print(f"{label}: {len(rows)} runs | foreign names in tool results: {len(seen)} | foreign reads: {len(reads)} | tools {out[label]['tools']}") + os.makedirs(os.path.join(ROOT, "runs", "analysis"), exist_ok=True) + json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "foreign_scan.json"), "w"), indent=1) + + +if __name__ == "__main__": + main() diff --git a/train/hf_train_bf16.py b/train/hf_train_bf16.py index c1a9063..a27d298 100644 --- a/train/hf_train_bf16.py +++ b/train/hf_train_bf16.py @@ -5,10 +5,10 @@ """Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go. Memory test first (a few dollars, finds the largest sequence length that fits, no data needed): - hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000 + hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000,64000 Real run (after the memory test and Kral's decision on the ratio): hf jobs uv run --flavor --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\ - --stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s1-epochs 1 --s2-epochs 3 --out erhankeseli/abap-mixed-adapter + --stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s2-epochs 3 --s2-loss-share 0.6 --out erhankeseli/abap-mixed-adapter Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns; none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut. @@ -35,6 +35,18 @@ def tokenize_masked(tok, text, spans, max_len): return ids, labels +def s1_epochs_for(share, s2_loss_tokens, s2_epochs, s1_tokens): + """Epochs of stage 1 so that stage 2 carries `share` of all loss tokens (Kral + Opus 2026-10-06: share = 0.6).""" + s2_total = s2_loss_tokens * s2_epochs + return (1 - share) / share * s2_total / max(s1_tokens, 1) + + +def copies(weight, rnd): + """Weight 1 = one copy per epoch, 2 = two, 0.5 = a copy in half of the epochs (a down-weight, not a drop).""" + w = max(float(weight), 0.0) + return int(w) + (1 if rnd.random() < w - int(w) else 0) + + def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed): rnd = random.Random(seed) rows = [] @@ -45,7 +57,7 @@ def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed): rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))] for ep in range(int(s2_epochs)): for r in stage2: - rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * max(1, round(float(r.get("weight", 1.0)))) + rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * copies(r.get("weight", 1.0), rnd) rnd.shuffle(rows) out, skipped = [], {"s1": 0, "s2": 0} for src, text, spans, _ in rows: @@ -65,8 +77,9 @@ def main(): ap.add_argument("--s1-file", default="train.jsonl") ap.add_argument("--s2-file", default="stage2_train.jsonl") ap.add_argument("--s2-valid", default="stage2_valid.jsonl") - ap.add_argument("--s1-epochs", type=float, default=1.0) + ap.add_argument("--s1-epochs", type=float, default=None, help="epochs of stage 1; default: computed from --s2-loss-share") ap.add_argument("--s2-epochs", type=float, default=3.0) + ap.add_argument("--s2-loss-share", type=float, default=0.6, help="share of the loss tokens that stage 2 carries (decision 2026-10-06: 0.6)") ap.add_argument("--max-seq", type=int, default=48000) ap.add_argument("--rank", type=int, default=16) ap.add_argument("--alpha", type=int, default=32) @@ -77,7 +90,7 @@ def main(): ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter") ap.add_argument("--save-every", type=int, default=0) ap.add_argument("--memory-test", action="store_true") - ap.add_argument("--sweep", default="16000,32000,48000", help="memory test: sequence lengths, tried in this order, stops at the first OOM") + ap.add_argument("--sweep", default="16000,32000,48000,64000", help="memory test: sequence lengths, tried in this order, stops at the first OOM") ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length") a = ap.parse_args() @@ -85,7 +98,8 @@ def main(): from unsloth import FastLanguageModel from huggingface_hub import HfApi, hf_hub_download token = os.environ["HF_TOKEN"] - model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=a.max_seq, load_in_4bit=False, dtype=torch.bfloat16, token=token) + seq_for_model = max([a.max_seq] + ([int(x) for x in a.sweep.split(",")] if a.memory_test else [])) + model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=seq_for_model, load_in_4bit=False, dtype=torch.bfloat16, token=token) model = FastLanguageModel.get_peft_model( model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none", target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"], @@ -132,6 +146,11 @@ def main(): p = hf_hub_download(repo, fname, repo_type="dataset", token=token) return [json.loads(l) for l in open(p)] s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file) + if a.s1_epochs is None: + s1_tok = sum(len(tok(r["text"], add_special_tokens=False)["input_ids"]) for r in s1) + s2_loss = sum(r["assistant_tokens"] * float(r.get("weight", 1.0)) for r in s2) + a.s1_epochs = s1_epochs_for(a.s2_loss_share, s2_loss, a.s2_epochs, s1_tok) + print("STAGE1 EPOCHS", round(a.s1_epochs, 3), "(stage 1 tokens", s1_tok, ", stage 2 loss tokens per epoch", round(s2_loss), ", share", a.s2_loss_share, ")", flush=True) ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006) loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex) print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens", diff --git a/train/hooks_example.py b/train/hooks_example.py index 580b544..262c530 100644 --- a/train/hooks_example.py +++ b/train/hooks_example.py @@ -10,17 +10,47 @@ def identity(row): return 1.0 +# Weights from the own-test mutation score (Kral + Opus 2026-10-06: down-weight, do not drop). Weight w means: w copies per epoch, a fraction is a +# copy in that share of the epochs (0.5 = every second epoch on average). The acceptance filter itself is unchanged. +OWN_TEST_WEIGHTS = { + "reliable, score >= 0.75": 1.0, # the own tests pass on the correct reference and kill at least 3 of 4 mutants + "reliable, score 0.5 to 0.75": 0.75, + "reliable, score < 0.5": 0.5, + "unreliable (own tests fail on the correct reference)": 0.5, # they may encode model specific behavior + "no own tests": 0.5, # accepted at 85 points at most; the habit of writing tests is part of the behavior we want + "no signal (PROG, no mutants, not scored)": 1.0, +} + + +def own_test_class(m): + if m is None: + return "no signal (PROG, no mutants, not scored)" + st = m.get("status") + if st == "no_own_tests": + return "no own tests" + if st != "scored": + return "no signal (PROG, no mutants, not scored)" + if not m.get("tests_pass_on_reference"): + return "unreliable (own tests fail on the correct reference)" + s = m.get("score") + if s is None: + return "no signal (PROG, no mutants, not scored)" + return "reliable, score >= 0.75" if s >= 0.75 else "reliable, score 0.5 to 0.75" if s >= 0.5 else "reliable, score < 0.5" + + def own_test_weight(row): - """Item D (own-test mutation score, metadata only for now): reads runs/traj//own_test_mutation.json when it exists, - stores it as extra data and does NOT drop or reweight (Kral + Opus 2026-10-06: do not change the acceptance yet).""" + """Item D: reads runs/traj//own_test_mutation.json, stores the score as extra data and sets the weight by OWN_TEST_WEIGHTS.""" run = row["id"].split("_r")[-1] + m = None for d in os.listdir(os.path.join(ROOT, "runs", "traj")): if d.startswith(run + "_"): p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json") if os.path.exists(p): m = json.load(open(p)) - return {"keep": True, "weight": 1.0, "extra": {"own_test_mutation": m.get("score"), "own_test_mutants": m.get("mutants")}} - return {"keep": True, "weight": 1.0} + break + cls = own_test_class(m) + return {"keep": True, "weight": OWN_TEST_WEIGHTS[cls], + "extra": {"own_test_class": cls, "own_test_mutation": (m or {}).get("score"), "own_test_reliable": bool((m or {}).get("tests_pass_on_reference"))}} def repair_up(row): diff --git a/train/mem_table.py b/train/mem_table.py new file mode 100644 index 0000000..f7f6947 --- /dev/null +++ b/train/mem_table.py @@ -0,0 +1,71 @@ +"""Writes docs/bf16-memory.md (estimates for the bf16 LoRA run of Qwen 3.8 27B by GPU and sequence length).""" +import json +import os + +ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +H, L, V, I = 5120, 64, 248320, 17408 +r = 16 +lora = 16 * r * ((H + 6144) + (H + 1024) + (H + 1024) + (6144 + H)) + 48 * r * ((H + 10240) + (H + 6144) + (6144 + H)) + L * r * ((H + I) * 2 + (I + H)) +W = 27.8e9 * 2 / 2**30 +lora_gb = lora * 10 / 2**30 + + +def est(seq, offload, chunked): + ckpt = 0.0 if offload else seq * H * 2 * L / 2**30 + layer = seq * I * 2 * 4 / 2**30 + 3 + logits = (seq * V * 2 * 3 / 2**30) if not chunked else 3.0 + return W + lora_gb + ckpt + layer + logits, dict(weights=W, lora=lora_gb, ckpt=ckpt, layer=layer, logits=logits) + + +gpus = [("a100-large", "1x A100 80 GB", 80), ("rtx-pro-6000", "1x RTX PRO 6000 96 GB", 96), ("h200", "1x H200 141 GB", 141), ("h200x2", "2x H200 282 GB (model parallel)", 282)] +tab = "| seq | checkpoints | loss | estimated peak GB | " + " | ".join(g[1] for g in gpus) + " |\n|---|---|---|---|" + "---|" * len(gpus) + "\n" +for seq in (16000, 32000, 48000, 64000): + for off in (True, False): + for ch in (True, False): + t, _ = est(seq, off, ch) + cells = ["fits" if t <= g[2] * 0.92 else ("tight" if t <= g[2] else "no") for g in gpus] + tab += f"| {seq // 1000}k | {'CPU offload (Unsloth)' if off else 'on the GPU'} | {'chunked' if ch else 'full logits'} | {t:.0f} | " + " | ".join(cells) + " |\n" +_, b48 = est(48000, True, True) +_, b64 = est(64000, True, True) +doc = f"""# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run) + +**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the +count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length, +stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000` +until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k). + +## Model (from config.json of Qwen3.8-27B) +64 layers (48 linear attention, 16 full attention), hidden {H}, MLP {I}, vocab {V}, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **{lora / 1e6:.0f} M trainable parameters**. + +## Memory by component (GB; estimate, not a measurement) +| component | 48k | 64k | how | +|---|---|---|---| +| weights bf16 | {b48['weights']:.1f} | {b64['weights']:.1f} | 27.8 B x 2 bytes | +| LoRA weights + grads + 8-bit Adam | {b48['lora']:.1f} | {b64['lora']:.1f} | {lora / 1e6:.0f} M x 10 bytes | +| layer inputs for the backward pass | {48000 * H * 2 * L / 2**30:.1f} on the GPU, about 0 with the Unsloth CPU offload (then {48000 * H * 2 * L / 2**30:.0f} GB host RAM) | {64000 * H * 2 * L / 2**30:.1f} / about 0 (host RAM {64000 * H * 2 * L / 2**30:.0f} GB) | seq x 5120 x 2 bytes x 64 layers | +| recompute peak of one layer | {b48['layer']:.1f} | {b64['layer']:.1f} | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) | +| logits and loss | {b48['logits']:.0f} chunked, {48000 * V * 2 * 3 / 2**30:.0f} full logits | {b64['logits']:.0f} chunked, {64000 * V * 2 * 3 / 2**30:.0f} full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) | + +## Fit by GPU (92 % of the card counted as usable) +{tab} +Reading: +- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy + (Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause. + Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k. +- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about {b48['weights'] + b48['lora'] + b48['layer'] + b48['logits']:.0f} GB at 48k, {b64['weights'] + b64['lora'] + b64['layer'] + b64['logits']:.0f} GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin; + **RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss. +- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test. + +## Order when the test is run +1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin. +2. If the loss is the problem: build the chunked loss (about 2 hours), repeat. +3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied). + +## Time and cost (estimate from the nf4 run: {6750 / 30.5:.0f} tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100) +Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total. +A100 about {3.3e6 / 300 / 3600:.1f} h (about {3.3e6 / 300 / 3600 * 2.5:.0f} USD), H200 about {3.3e6 / 650 / 3600:.1f} h (about {3.3e6 / 650 / 3600 * 5:.0f} USD). +With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens, +H200 about {10e6 / 650 / 3600:.1f} h (about {10e6 / 650 / 3600 * 5:.0f} USD), A100 about {10e6 / 300 / 3600:.0f} h (about {10e6 / 300 / 3600 * 2.5:.0f} USD). The real number comes from the memory test (`step_seconds`). +""" +open(os.path.join(ROOT, "docs", "bf16-memory.md"), "w").write(doc) +print("written") diff --git a/train/requeue_foreign.py b/train/requeue_foreign.py new file mode 100644 index 0000000..d02d790 --- /dev/null +++ b/train/requeue_foreign.py @@ -0,0 +1,75 @@ +"""Trajectories in which the model read another run's leftover object go back to the pending pool (Kral + Opus 2026-10-06). + + python3 train/requeue_foreign.py dry run + python3 train/requeue_foreign.py --apply moves their rows from runs/traj/summary.jsonl to runs/traj/summary_excluded.jsonl + (backup: summary.jsonl.bak-