Compare commits
2 Commits
4002d889d5
...
8ef85c2713
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8ef85c2713 | ||
|
|
a465c33e3f |
@@ -1,21 +1,24 @@
|
||||
# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
|
||||
|
||||
**No GPU job is started without Kral's go.** Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000` is the memory test, no data needed, 3 optimizer steps per length, stops at the first OOM, uploads the result as `memtest_*.json` to the output repo).
|
||||
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
|
||||
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
|
||||
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
|
||||
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
|
||||
|
||||
## Model (from config.json of Qwen3.8-27B)
|
||||
64 layers (48 linear attention, 16 full attention), hidden 5120, MLP 17408, vocab 248320, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **107 M trainable parameters**.
|
||||
|
||||
## Memory by component at 48k tokens (GB; the estimate, not a measurement)
|
||||
| component | GB | how |
|
||||
|---|---|---|
|
||||
| weights bf16 | 51.8 | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | 1.0 | 107 M x 10 bytes |
|
||||
| layer inputs for the backward pass (checkpoints) | 29.3 on the GPU, about 0 with the Unsloth CPU offload | 48k x 5120 x 2 bytes x 64 layers; the offload needs 29 GB of host RAM (all flavors have 142 GB or more) |
|
||||
| recompute peak of one layer | 9.2 | MLP tensors 48k x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | 3 chunked, 67 with full logits | full logits: 48k x 248320 x (bf16 + fp32 upcast + grad); only a chunked or fused cross entropy is realistic |
|
||||
## Memory by component (GB; estimate, not a measurement)
|
||||
| component | 48k | 64k | how |
|
||||
|---|---|---|---|
|
||||
| weights bf16 | 51.8 | 51.8 | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | 1.0 | 1.0 | 107 M x 10 bytes |
|
||||
| layer inputs for the backward pass | 29.3 on the GPU, about 0 with the Unsloth CPU offload (then 29 GB host RAM) | 39.1 / about 0 (host RAM 39 GB) | seq x 5120 x 2 bytes x 64 layers |
|
||||
| recompute peak of one layer | 9.2 | 11.3 | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | 3 chunked, 67 full logits | 3 chunked, 89 full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
|
||||
|
||||
## Fit by GPU (92 % of the card counted as usable)
|
||||
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (needs model parallel) |
|
||||
| seq | checkpoints | loss | estimated peak GB | 1x A100 80 GB | 1x RTX PRO 6000 96 GB | 1x H200 141 GB | 2x H200 282 GB (model parallel) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 16k | CPU offload (Unsloth) | chunked | 61 | fits | fits | fits | fits |
|
||||
| 16k | CPU offload (Unsloth) | full logits | 80 | no | fits | fits | fits |
|
||||
@@ -29,20 +32,26 @@
|
||||
| 48k | CPU offload (Unsloth) | full logits | 129 | no | no | fits | fits |
|
||||
| 48k | on the GPU | chunked | 94 | no | tight | fits | fits |
|
||||
| 48k | on the GPU | full logits | 158 | no | no | no | fits |
|
||||
| 64k | CPU offload (Unsloth) | chunked | 67 | fits | fits | fits | fits |
|
||||
| 64k | CPU offload (Unsloth) | full logits | 153 | no | no | no | fits |
|
||||
| 64k | on the GPU | chunked | 106 | no | no | fits | fits |
|
||||
| 64k | on the GPU | full logits | 192 | no | no | no | fits |
|
||||
|
||||
Reading:
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, with the CPU offload) fits 32k and 48k; the A100 and the RTX PRO 6000 do not. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is not known). The memory test shows it at once: if `seq_len 16000` already fails on an H200, the loss is the cause.
|
||||
Fallback (not built yet, small): compute the hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour) is only possible with the CPU offload and a chunked loss** (about 65 GB estimated at 48k: no margin). H200 141 GB (5 USD/hour) fits with margin;
|
||||
RTX PRO 6000 96 GB (2.75 USD/hour) fits with the offload and the chunked loss (est. 65 GB).
|
||||
- Samples over 32k are a minority (p90 39560 tokens in the current data, p50 24.8k): if 48k does not fit, 32k would drop 15 of 52 samples (29 %; 10 of them DDLS), so test 48k first.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): possible only with model parallelism (`device_map`), one card works at a time, Unsloth multi-GPU is limited; not recommended before the single card test.
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
|
||||
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about 65 GB at 48k, 67 GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
|
||||
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
|
||||
|
||||
## Proposed order
|
||||
1. Memory test on **h200** (about 30 minutes, 2.5 USD): sweep 16000,32000,48000 with the offload. If 48k fits with margin: stay on H200 or try rtx-pro-6000 for the real run.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours of work), repeat the test.
|
||||
3. Real run after the stage 1 : stage 2 ratio decision (below).
|
||||
## Order when the test is run
|
||||
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
|
||||
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
|
||||
|
||||
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 is faster, H200 about 2 to 2.5 times an A100)
|
||||
Tokens of the proposed run (1 epoch stage 1 and 3 epochs of the 47 stage 2 samples as built today: 2.53 M + 3 x 1.28 M = 6.4 M tokens): A100 about 5.9 h (about 15 USD), H200 about 2.7 h (about 14 USD). The real number comes from the memory test (`step_seconds`).
|
||||
## Time and cost (estimate from the nf4 run: 221 tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
|
||||
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
|
||||
A100 about 3.1 h (about 8 USD), H200 about 1.4 h (about 7 USD).
|
||||
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
|
||||
H200 about 4.3 h (about 21 USD), A100 about 9 h (about 23 USD). The real number comes from the memory test (`step_seconds`).
|
||||
|
||||
29
docs/foreign-objects-report.md
Normal file
29
docs/foreign-objects-report.md
Normal file
@@ -0,0 +1,29 @@
|
||||
# Other runs' objects in tool results: which runs are affected (2026-10-06)
|
||||
|
||||
Cause: the model sometimes names its own helper or test class `ZCL_<run prefix>_...` (prefix inside the name). The teardown looked only for names that start with the prefix, so such
|
||||
classes stayed in A4H and showed up in `sap_inactive_objects`, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, `harness/sweep.py`).
|
||||
Scan: `train/foreign_scan.py` (read only; result `runs/analysis/foreign_scan.json`). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object.
|
||||
|
||||
| group | runs | with foreign names in a tool result | with a foreign read | tools |
|
||||
|---|---|---|---|---|
|
||||
| official baseline Qwen (MacBook, 11 tasks) | 11 | 0 | 0 | - |
|
||||
| Devstral baseline (dropped) | 11 | 0 | 0 | - |
|
||||
| older Qwen baselines (archive) | 8 | 0 | 0 | - |
|
||||
| smoke | 1 | 0 | 0 | - |
|
||||
| eval empirical filter (DeepSeek on eval candidates) | 178 | 25 | 4 | sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5 |
|
||||
| eval generation validation (oracle, null, mutants) | 1344 | 0 | 0 | - |
|
||||
| early pilot runs | 16 | 1 | 0 | sap_inactive_objects 1 |
|
||||
| training trajectories (DeepSeek) | 117 | 67 | 15 | sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1 |
|
||||
| series A (local Qwen) | 20 | 3 | 0 | sap_inactive_objects 2, sap_search_object 1 |
|
||||
|
||||
## Reading
|
||||
- **Official baseline (Qwen, 11 tasks, 2026-10-04): not affected** (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers,
|
||||
and none of the 11 baseline runs called `sap_inactive_objects` (checked: 0 calls). **The baseline numbers stay as they are. Nothing was changed.**
|
||||
- **Eval generation (1344 validation runs: oracle, null, mutants): not affected** (they call no list or search tool).
|
||||
- **Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one**
|
||||
(G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there
|
||||
was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs).
|
||||
- **Early pilot runs:** 1 of 16 saw a foreign name in an inactive list.
|
||||
- **Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one.** Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool
|
||||
(`train/requeue_foreign.py`, rows in `runs/traj/summary_excluded.jsonl`) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder.
|
||||
- **Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads.** The scores (3 of 20) are not affected by a read; the series ran before the fix.
|
||||
20
docs/own-test-mutation.md
Normal file
20
docs/own-test-mutation.md
Normal file
@@ -0,0 +1,20 @@
|
||||
# Own-test mutation score of the accepted trajectories (item D, metadata only)
|
||||
|
||||
Module `harness/owntests.py`, results `runs/traj/<run>/own_test_mutation.json` (not in git), report by `train/own_test_report.py`.
|
||||
The model's own unit tests (testclasses include or global test classes) are run against the faulty references of the task (`faulty/`, mutants that the hidden tests kill).
|
||||
The correct reference must pass the model's tests first (otherwise the tests encode model specific behavior). score = killed / (killed + survived).
|
||||
**Metadata only: the acceptance filter is not changed.** The builder takes the score through `--hook hooks_example:own_test_weight` (field `own_test_mutation`).
|
||||
|
||||
## Result
|
||||
90 accepted trajectories looked at: {'scored': 63, 'not_supported': 5, 'no_own_tests': 22}.
|
||||
- Scored: 63; the model's tests also passed on the correct reference: 53 (the others are not reliable: a test that fails on the correct solution kills every mutant).
|
||||
- Mean score over the reliable ones: 0.99; median 1.00; distribution {'1.0': 50, '>=0.75': 3} (n = 53).
|
||||
- By kind (mean, n): {CLAS: 1.00 (45), DDLS: 0.90 (5), FUNC: 1.00 (3)}
|
||||
- `no_own_tests`: 22 trajectories were accepted without any own test (at most 85 points); `no_mutants`: 0; `not_supported` (PROG: tests are inside the program): 5.
|
||||
|
||||
## Weakest (score under 0.75, tests pass on the reference)
|
||||
|
||||
|
||||
## Use
|
||||
Not used for filtering yet. Candidates for a later rule (Kral + Opus decide): drop or down-weight trajectories with a reliable score under 0.5 or with `tests_pass_on_reference` false;
|
||||
prefer trajectories without own tests last. The hook example shows where the weight goes (`train/hooks_example.py`).
|
||||
@@ -63,6 +63,15 @@ Second attempts are part of the 130 runs. If the new reset gives 60 usage, this
|
||||
- Tool budget 100 for CDS tasks (eval stays 60). 20 tool schemas in every sample. Token note: p95 48k, max 72k.
|
||||
- Summary every 50 accepted trajectories in `train/STATE.md` and `docs/yol-haritasi.md`, with a commit.
|
||||
|
||||
## 5b. Check on 11 October (the last work before the reset; Claude does it when Kral asks)
|
||||
1. `git status` clean and pushed; `python3 -m harness.restart_plan --panel <value>` prints the numbers (panel value from Kral, expected 0 after the reset).
|
||||
2. Services: A4H up, MCP answers, `python3 scripts_probe/lockprobe.py` shows no stale lock (only ATC runtime and debugger listener entries), `python3 -m harness.sweep` shows no leftover.
|
||||
3. `.env`: `BUDGET_CYCLE_START=2026-10-12`, `BUDGET_LIMIT_USD`, `BUDGET_RESERVE_USD=8`; no `STOP` flag, no `STOPPED.txt`; no running controller (`runs/pipeline/controller.lock`).
|
||||
4. Order of the first hour: step 0 eval slots for the new kinds (`evalset plan-new`, `run-new`, `run-new-k`), then the pending tasks of the kinds below target. The 9 trajectories that were
|
||||
moved back (`runs/traj/summary_excluded.jsonl`: FUNC 3, DDLS 5, TABL 1) are pending again and run in the normal order.
|
||||
5. Settings to confirm: workers 2, `STREAM_GUARD` unset for the first runs and then 9000 for 5 DDLS/PROG runs (`docs/empty-response.md`), tool budget 100 for CDS.
|
||||
6. Not started and not to be started before Kral's go: the bf16 memory test (wait until the data is near the size of the first SFT run), second A4H, multi-host.
|
||||
|
||||
## 6. After the restart
|
||||
|
||||
Open decisions: second A4H (not started, multi-host code later), bf16 memory test at 48k, the second teacher test (the
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Stage 2 training set, build of 2026-10-06 06:15:28
|
||||
# Stage 2 training set, build of 2026-10-06 09:03:41
|
||||
|
||||
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
||||
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||
@@ -8,9 +8,7 @@ HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectorie
|
||||
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
|
||||
Dropped:
|
||||
- FUNC: {'read_of_another_runs_object': 3}
|
||||
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
|
||||
- TABL: {'read_of_another_runs_object': 1}
|
||||
- DDLS: {'over_48k': 5}
|
||||
- INTF: {'over_48k': 1}
|
||||
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
||||
|
||||
@@ -34,7 +32,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
|
||||
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
|
||||
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
@@ -52,7 +50,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
|
||||
|
||||
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
|
||||
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
|
||||
|
||||
| class of the trajectory | weight | reason |
|
||||
|---|---|---|
|
||||
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
|
||||
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
|
||||
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
|
||||
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
|
||||
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
|
||||
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
|
||||
|
||||
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
|
||||
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
|
||||
|
||||
## Other decisions of 2026-10-06
|
||||
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
|
||||
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
|
||||
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
|
||||
- **Memory test:** waits until the data is near the size of the first SFT run.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
|
||||
@@ -43,7 +43,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
|
||||
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
|
||||
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
@@ -61,7 +61,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
|
||||
|
||||
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
|
||||
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
|
||||
|
||||
| class of the trajectory | weight | reason |
|
||||
|---|---|---|
|
||||
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
|
||||
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
|
||||
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
|
||||
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
|
||||
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
|
||||
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
|
||||
|
||||
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
|
||||
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
|
||||
|
||||
## Other decisions of 2026-10-06
|
||||
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
|
||||
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
|
||||
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
|
||||
- **Memory test:** waits until the data is near the size of the first SFT run.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
|
||||
79
train/foreign_scan.py
Normal file
79
train/foreign_scan.py
Normal file
@@ -0,0 +1,79 @@
|
||||
"""Read-only scan: which runs saw other runs' objects in tool results, or read them (teardown/proxy bug of 2026-10-06).
|
||||
|
||||
python3 train/foreign_scan.py -> prints a summary and writes runs/analysis/foreign_scan.json
|
||||
"""
|
||||
import glob
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from collections import Counter
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
sys.path.insert(0, ROOT)
|
||||
from harness.task import prefix_for # noqa: E402
|
||||
|
||||
START = re.compile(r"Z\d[0-9A-Z]{6}_")
|
||||
MID_NAME = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
|
||||
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
|
||||
"sap_syntax_check", "sap_atc_run")
|
||||
GROUPS = [("official baseline Qwen (MacBook, 11 tasks)", "runs/stage1/baseline_qwen/*"),
|
||||
("Devstral baseline (dropped)", "runs/archive/devstral/baseline_devstral/*"),
|
||||
("older Qwen baselines (archive)", "runs/archive/qwen38/*/*"),
|
||||
("smoke", "runs/stage1/smoke/*"),
|
||||
("eval empirical filter (DeepSeek on eval candidates)", "runs/emp/*"),
|
||||
("eval generation validation (oracle, null, mutants)", "runs/gen/*"),
|
||||
("early pilot runs", "runs/2*_*"),
|
||||
("training trajectories (DeepSeek)", "runs/traj/*_G*"),
|
||||
("series A (local Qwen)", "runs/local_qwen/runs/*")]
|
||||
|
||||
|
||||
def own_prefix(d):
|
||||
m = re.match(r"^(\d+)_([GT]\d+)_", os.path.basename(d))
|
||||
return prefix_for(int(m.group(1)), m.group(2)).upper() if m else None
|
||||
|
||||
|
||||
def scan(d):
|
||||
own = own_prefix(d)
|
||||
p = os.path.join(d, "trajectory.jsonl")
|
||||
if not own or not os.path.exists(p):
|
||||
return None
|
||||
seen, reads, tools = set(), 0, Counter()
|
||||
for l in open(p):
|
||||
try:
|
||||
e = json.loads(l)
|
||||
except ValueError:
|
||||
continue
|
||||
if "tool" not in e:
|
||||
continue
|
||||
text = (e.get("result") or "").upper()
|
||||
found = {x for x in START.findall(text) if x != own and x[1].isdigit()}
|
||||
found |= {m.group(1).upper() for m in re.finditer(r"\b[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", text) if m.group(1).upper() != own}
|
||||
if found:
|
||||
seen |= found
|
||||
tools[e["tool"]] += 1
|
||||
name = str((e.get("args") or {}).get("objectName", "")).upper()
|
||||
m = START.match(name) or MID_NAME.match(name)
|
||||
pre = (m.group(1) if m and m.re is MID_NAME else (m.group(0) if m else "")).upper()
|
||||
if e["tool"] in READ and pre and pre != own:
|
||||
reads += 1
|
||||
return {"run": os.path.basename(d), "foreign_prefixes_seen": len(seen), "foreign_reads": reads, "tools": dict(tools)}
|
||||
|
||||
|
||||
def main():
|
||||
out = {}
|
||||
for label, pat in GROUPS:
|
||||
dirs = [d for d in sorted(glob.glob(os.path.join(ROOT, pat))) if os.path.isdir(d)]
|
||||
rows = [r for r in (scan(d) for d in dirs) if r]
|
||||
seen = [r for r in rows if r["foreign_prefixes_seen"]]
|
||||
reads = [r for r in rows if r["foreign_reads"]]
|
||||
out[label] = {"runs": len(rows), "runs_with_foreign_names": len(seen), "runs_with_foreign_reads": len(reads),
|
||||
"tools": dict(sum((Counter(r["tools"]) for r in seen), Counter())),
|
||||
"affected": [r["run"] for r in seen][:40], "read_runs": [r["run"] for r in reads]}
|
||||
print(f"{label}: {len(rows)} runs | foreign names in tool results: {len(seen)} | foreign reads: {len(reads)} | tools {out[label]['tools']}")
|
||||
os.makedirs(os.path.join(ROOT, "runs", "analysis"), exist_ok=True)
|
||||
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "foreign_scan.json"), "w"), indent=1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -5,10 +5,10 @@
|
||||
"""Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go.
|
||||
|
||||
Memory test first (a few dollars, finds the largest sequence length that fits, no data needed):
|
||||
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000
|
||||
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000,64000
|
||||
Real run (after the memory test and Kral's decision on the ratio):
|
||||
hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\
|
||||
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s1-epochs 1 --s2-epochs 3 --out erhankeseli/abap-mixed-adapter
|
||||
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s2-epochs 3 --s2-loss-share 0.6 --out erhankeseli/abap-mixed-adapter
|
||||
|
||||
Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns;
|
||||
none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut.
|
||||
@@ -35,6 +35,18 @@ def tokenize_masked(tok, text, spans, max_len):
|
||||
return ids, labels
|
||||
|
||||
|
||||
def s1_epochs_for(share, s2_loss_tokens, s2_epochs, s1_tokens):
|
||||
"""Epochs of stage 1 so that stage 2 carries `share` of all loss tokens (Kral + Opus 2026-10-06: share = 0.6)."""
|
||||
s2_total = s2_loss_tokens * s2_epochs
|
||||
return (1 - share) / share * s2_total / max(s1_tokens, 1)
|
||||
|
||||
|
||||
def copies(weight, rnd):
|
||||
"""Weight 1 = one copy per epoch, 2 = two, 0.5 = a copy in half of the epochs (a down-weight, not a drop)."""
|
||||
w = max(float(weight), 0.0)
|
||||
return int(w) + (1 if rnd.random() < w - int(w) else 0)
|
||||
|
||||
|
||||
def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
|
||||
rnd = random.Random(seed)
|
||||
rows = []
|
||||
@@ -45,7 +57,7 @@ def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
|
||||
rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))]
|
||||
for ep in range(int(s2_epochs)):
|
||||
for r in stage2:
|
||||
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * max(1, round(float(r.get("weight", 1.0))))
|
||||
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * copies(r.get("weight", 1.0), rnd)
|
||||
rnd.shuffle(rows)
|
||||
out, skipped = [], {"s1": 0, "s2": 0}
|
||||
for src, text, spans, _ in rows:
|
||||
@@ -65,8 +77,9 @@ def main():
|
||||
ap.add_argument("--s1-file", default="train.jsonl")
|
||||
ap.add_argument("--s2-file", default="stage2_train.jsonl")
|
||||
ap.add_argument("--s2-valid", default="stage2_valid.jsonl")
|
||||
ap.add_argument("--s1-epochs", type=float, default=1.0)
|
||||
ap.add_argument("--s1-epochs", type=float, default=None, help="epochs of stage 1; default: computed from --s2-loss-share")
|
||||
ap.add_argument("--s2-epochs", type=float, default=3.0)
|
||||
ap.add_argument("--s2-loss-share", type=float, default=0.6, help="share of the loss tokens that stage 2 carries (decision 2026-10-06: 0.6)")
|
||||
ap.add_argument("--max-seq", type=int, default=48000)
|
||||
ap.add_argument("--rank", type=int, default=16)
|
||||
ap.add_argument("--alpha", type=int, default=32)
|
||||
@@ -77,7 +90,7 @@ def main():
|
||||
ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter")
|
||||
ap.add_argument("--save-every", type=int, default=0)
|
||||
ap.add_argument("--memory-test", action="store_true")
|
||||
ap.add_argument("--sweep", default="16000,32000,48000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
|
||||
ap.add_argument("--sweep", default="16000,32000,48000,64000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
|
||||
ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length")
|
||||
a = ap.parse_args()
|
||||
|
||||
@@ -85,7 +98,8 @@ def main():
|
||||
from unsloth import FastLanguageModel
|
||||
from huggingface_hub import HfApi, hf_hub_download
|
||||
token = os.environ["HF_TOKEN"]
|
||||
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=a.max_seq, load_in_4bit=False, dtype=torch.bfloat16, token=token)
|
||||
seq_for_model = max([a.max_seq] + ([int(x) for x in a.sweep.split(",")] if a.memory_test else []))
|
||||
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=seq_for_model, load_in_4bit=False, dtype=torch.bfloat16, token=token)
|
||||
model = FastLanguageModel.get_peft_model(
|
||||
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
|
||||
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"],
|
||||
@@ -132,6 +146,11 @@ def main():
|
||||
p = hf_hub_download(repo, fname, repo_type="dataset", token=token)
|
||||
return [json.loads(l) for l in open(p)]
|
||||
s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file)
|
||||
if a.s1_epochs is None:
|
||||
s1_tok = sum(len(tok(r["text"], add_special_tokens=False)["input_ids"]) for r in s1)
|
||||
s2_loss = sum(r["assistant_tokens"] * float(r.get("weight", 1.0)) for r in s2)
|
||||
a.s1_epochs = s1_epochs_for(a.s2_loss_share, s2_loss, a.s2_epochs, s1_tok)
|
||||
print("STAGE1 EPOCHS", round(a.s1_epochs, 3), "(stage 1 tokens", s1_tok, ", stage 2 loss tokens per epoch", round(s2_loss), ", share", a.s2_loss_share, ")", flush=True)
|
||||
ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006)
|
||||
loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex)
|
||||
print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens",
|
||||
|
||||
@@ -10,17 +10,47 @@ def identity(row):
|
||||
return 1.0
|
||||
|
||||
|
||||
# Weights from the own-test mutation score (Kral + Opus 2026-10-06: down-weight, do not drop). Weight w means: w copies per epoch, a fraction is a
|
||||
# copy in that share of the epochs (0.5 = every second epoch on average). The acceptance filter itself is unchanged.
|
||||
OWN_TEST_WEIGHTS = {
|
||||
"reliable, score >= 0.75": 1.0, # the own tests pass on the correct reference and kill at least 3 of 4 mutants
|
||||
"reliable, score 0.5 to 0.75": 0.75,
|
||||
"reliable, score < 0.5": 0.5,
|
||||
"unreliable (own tests fail on the correct reference)": 0.5, # they may encode model specific behavior
|
||||
"no own tests": 0.5, # accepted at 85 points at most; the habit of writing tests is part of the behavior we want
|
||||
"no signal (PROG, no mutants, not scored)": 1.0,
|
||||
}
|
||||
|
||||
|
||||
def own_test_class(m):
|
||||
if m is None:
|
||||
return "no signal (PROG, no mutants, not scored)"
|
||||
st = m.get("status")
|
||||
if st == "no_own_tests":
|
||||
return "no own tests"
|
||||
if st != "scored":
|
||||
return "no signal (PROG, no mutants, not scored)"
|
||||
if not m.get("tests_pass_on_reference"):
|
||||
return "unreliable (own tests fail on the correct reference)"
|
||||
s = m.get("score")
|
||||
if s is None:
|
||||
return "no signal (PROG, no mutants, not scored)"
|
||||
return "reliable, score >= 0.75" if s >= 0.75 else "reliable, score 0.5 to 0.75" if s >= 0.5 else "reliable, score < 0.5"
|
||||
|
||||
|
||||
def own_test_weight(row):
|
||||
"""Item D (own-test mutation score, metadata only for now): reads runs/traj/<run>/own_test_mutation.json when it exists,
|
||||
stores it as extra data and does NOT drop or reweight (Kral + Opus 2026-10-06: do not change the acceptance yet)."""
|
||||
"""Item D: reads runs/traj/<run>/own_test_mutation.json, stores the score as extra data and sets the weight by OWN_TEST_WEIGHTS."""
|
||||
run = row["id"].split("_r")[-1]
|
||||
m = None
|
||||
for d in os.listdir(os.path.join(ROOT, "runs", "traj")):
|
||||
if d.startswith(run + "_"):
|
||||
p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json")
|
||||
if os.path.exists(p):
|
||||
m = json.load(open(p))
|
||||
return {"keep": True, "weight": 1.0, "extra": {"own_test_mutation": m.get("score"), "own_test_mutants": m.get("mutants")}}
|
||||
return {"keep": True, "weight": 1.0}
|
||||
break
|
||||
cls = own_test_class(m)
|
||||
return {"keep": True, "weight": OWN_TEST_WEIGHTS[cls],
|
||||
"extra": {"own_test_class": cls, "own_test_mutation": (m or {}).get("score"), "own_test_reliable": bool((m or {}).get("tests_pass_on_reference"))}}
|
||||
|
||||
|
||||
def repair_up(row):
|
||||
|
||||
71
train/mem_table.py
Normal file
71
train/mem_table.py
Normal file
@@ -0,0 +1,71 @@
|
||||
"""Writes docs/bf16-memory.md (estimates for the bf16 LoRA run of Qwen 3.8 27B by GPU and sequence length)."""
|
||||
import json
|
||||
import os
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
H, L, V, I = 5120, 64, 248320, 17408
|
||||
r = 16
|
||||
lora = 16 * r * ((H + 6144) + (H + 1024) + (H + 1024) + (6144 + H)) + 48 * r * ((H + 10240) + (H + 6144) + (6144 + H)) + L * r * ((H + I) * 2 + (I + H))
|
||||
W = 27.8e9 * 2 / 2**30
|
||||
lora_gb = lora * 10 / 2**30
|
||||
|
||||
|
||||
def est(seq, offload, chunked):
|
||||
ckpt = 0.0 if offload else seq * H * 2 * L / 2**30
|
||||
layer = seq * I * 2 * 4 / 2**30 + 3
|
||||
logits = (seq * V * 2 * 3 / 2**30) if not chunked else 3.0
|
||||
return W + lora_gb + ckpt + layer + logits, dict(weights=W, lora=lora_gb, ckpt=ckpt, layer=layer, logits=logits)
|
||||
|
||||
|
||||
gpus = [("a100-large", "1x A100 80 GB", 80), ("rtx-pro-6000", "1x RTX PRO 6000 96 GB", 96), ("h200", "1x H200 141 GB", 141), ("h200x2", "2x H200 282 GB (model parallel)", 282)]
|
||||
tab = "| seq | checkpoints | loss | estimated peak GB | " + " | ".join(g[1] for g in gpus) + " |\n|---|---|---|---|" + "---|" * len(gpus) + "\n"
|
||||
for seq in (16000, 32000, 48000, 64000):
|
||||
for off in (True, False):
|
||||
for ch in (True, False):
|
||||
t, _ = est(seq, off, ch)
|
||||
cells = ["fits" if t <= g[2] * 0.92 else ("tight" if t <= g[2] else "no") for g in gpus]
|
||||
tab += f"| {seq // 1000}k | {'CPU offload (Unsloth)' if off else 'on the GPU'} | {'chunked' if ch else 'full logits'} | {t:.0f} | " + " | ".join(cells) + " |\n"
|
||||
_, b48 = est(48000, True, True)
|
||||
_, b64 = est(64000, True, True)
|
||||
doc = f"""# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
|
||||
|
||||
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
|
||||
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
|
||||
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
|
||||
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
|
||||
|
||||
## Model (from config.json of Qwen3.8-27B)
|
||||
64 layers (48 linear attention, 16 full attention), hidden {H}, MLP {I}, vocab {V}, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **{lora / 1e6:.0f} M trainable parameters**.
|
||||
|
||||
## Memory by component (GB; estimate, not a measurement)
|
||||
| component | 48k | 64k | how |
|
||||
|---|---|---|---|
|
||||
| weights bf16 | {b48['weights']:.1f} | {b64['weights']:.1f} | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | {b48['lora']:.1f} | {b64['lora']:.1f} | {lora / 1e6:.0f} M x 10 bytes |
|
||||
| layer inputs for the backward pass | {48000 * H * 2 * L / 2**30:.1f} on the GPU, about 0 with the Unsloth CPU offload (then {48000 * H * 2 * L / 2**30:.0f} GB host RAM) | {64000 * H * 2 * L / 2**30:.1f} / about 0 (host RAM {64000 * H * 2 * L / 2**30:.0f} GB) | seq x 5120 x 2 bytes x 64 layers |
|
||||
| recompute peak of one layer | {b48['layer']:.1f} | {b64['layer']:.1f} | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | {b48['logits']:.0f} chunked, {48000 * V * 2 * 3 / 2**30:.0f} full logits | {b64['logits']:.0f} chunked, {64000 * V * 2 * 3 / 2**30:.0f} full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
|
||||
|
||||
## Fit by GPU (92 % of the card counted as usable)
|
||||
{tab}
|
||||
Reading:
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
|
||||
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about {b48['weights'] + b48['lora'] + b48['layer'] + b48['logits']:.0f} GB at 48k, {b64['weights'] + b64['lora'] + b64['layer'] + b64['logits']:.0f} GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
|
||||
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
|
||||
|
||||
## Order when the test is run
|
||||
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
|
||||
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
|
||||
|
||||
## Time and cost (estimate from the nf4 run: {6750 / 30.5:.0f} tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
|
||||
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
|
||||
A100 about {3.3e6 / 300 / 3600:.1f} h (about {3.3e6 / 300 / 3600 * 2.5:.0f} USD), H200 about {3.3e6 / 650 / 3600:.1f} h (about {3.3e6 / 650 / 3600 * 5:.0f} USD).
|
||||
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
|
||||
H200 about {10e6 / 650 / 3600:.1f} h (about {10e6 / 650 / 3600 * 5:.0f} USD), A100 about {10e6 / 300 / 3600:.0f} h (about {10e6 / 300 / 3600 * 2.5:.0f} USD). The real number comes from the memory test (`step_seconds`).
|
||||
"""
|
||||
open(os.path.join(ROOT, "docs", "bf16-memory.md"), "w").write(doc)
|
||||
print("written")
|
||||
75
train/requeue_foreign.py
Normal file
75
train/requeue_foreign.py
Normal file
@@ -0,0 +1,75 @@
|
||||
"""Trajectories in which the model read another run's leftover object go back to the pending pool (Kral + Opus 2026-10-06).
|
||||
|
||||
python3 train/requeue_foreign.py dry run
|
||||
python3 train/requeue_foreign.py --apply moves their rows from runs/traj/summary.jsonl to runs/traj/summary_excluded.jsonl
|
||||
(backup: summary.jsonl.bak-<time>); the task then counts as not run and is run again
|
||||
after the reset in the normal order (kind deficit). The run folders stay.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sys
|
||||
import time
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
sys.path.insert(0, ROOT)
|
||||
sys.path.insert(0, os.path.join(ROOT, "train"))
|
||||
import accept as acc # noqa: E402
|
||||
from harness import mix # noqa: E402
|
||||
|
||||
START = __import__("re").compile(r"^(Z\d[0-9A-Z]{6}_)", __import__("re").I)
|
||||
MID = __import__("re").compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", __import__("re").I)
|
||||
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
|
||||
"sap_syntax_check", "sap_atc_run")
|
||||
|
||||
|
||||
def foreign_reads(rec):
|
||||
own = rec["prefix"].upper()
|
||||
n = 0
|
||||
for m in rec["messages"]:
|
||||
if m["role"] != "assistant":
|
||||
continue
|
||||
for c in m.get("tool_calls") or []:
|
||||
if c["function"]["name"] not in READ:
|
||||
continue
|
||||
a = c["function"].get("arguments") or "{}"
|
||||
try:
|
||||
a = json.loads(a) if isinstance(a, str) else a
|
||||
except ValueError:
|
||||
a = {}
|
||||
name = str(a.get("objectName", "")).upper()
|
||||
mm = START.match(name) or MID.match(name)
|
||||
if mm and mm.group(1).upper() != own:
|
||||
n += 1
|
||||
return n
|
||||
|
||||
|
||||
def main():
|
||||
apply = "--apply" in sys.argv
|
||||
path = os.path.join(ROOT, "runs", "traj", "summary.jsonl")
|
||||
rows = [json.loads(l) for l in open(path)]
|
||||
keep, out = [], []
|
||||
for r in rows:
|
||||
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
|
||||
if os.path.exists(p):
|
||||
rec = json.load(open(p))
|
||||
if acc.judge(rec, r, 80)[0]:
|
||||
n = foreign_reads(rec)
|
||||
if n:
|
||||
out.append(dict(r, excluded_reason="read of another run's object", foreign_reads=n, kind=mix.kind_of_task_dir(r["task"]),
|
||||
excluded_at=time.strftime("%F %T")))
|
||||
continue
|
||||
keep.append(r)
|
||||
print(len(out), "accepted trajectories with a foreign read:", [(o["task"], o["attempt"], o["kind"], o["foreign_reads"]) for o in out])
|
||||
if not apply:
|
||||
return
|
||||
shutil.copy(path, path + ".bak-" + time.strftime("%Y%m%d-%H%M%S"))
|
||||
with open(os.path.join(ROOT, "runs", "traj", "summary_excluded.jsonl"), "a") as f:
|
||||
for o in out:
|
||||
f.write(json.dumps(o) + "\n")
|
||||
open(path, "w").write("".join(json.dumps(r) + "\n" for r in keep))
|
||||
print("moved; summary.jsonl now", len(keep), "rows")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user