Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -43,7 +43,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
|
||||
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
|
||||
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
@@ -61,7 +61,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
|
||||
|
||||
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
|
||||
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
|
||||
|
||||
| class of the trajectory | weight | reason |
|
||||
|---|---|---|
|
||||
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
|
||||
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
|
||||
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
|
||||
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
|
||||
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
|
||||
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
|
||||
|
||||
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
|
||||
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
|
||||
|
||||
## Other decisions of 2026-10-06
|
||||
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
|
||||
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
|
||||
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
|
||||
- **Memory test:** waits until the data is near the size of the first SFT run.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
|
||||
79
train/foreign_scan.py
Normal file
79
train/foreign_scan.py
Normal file
@@ -0,0 +1,79 @@
|
||||
"""Read-only scan: which runs saw other runs' objects in tool results, or read them (teardown/proxy bug of 2026-10-06).
|
||||
|
||||
python3 train/foreign_scan.py -> prints a summary and writes runs/analysis/foreign_scan.json
|
||||
"""
|
||||
import glob
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from collections import Counter
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
sys.path.insert(0, ROOT)
|
||||
from harness.task import prefix_for # noqa: E402
|
||||
|
||||
START = re.compile(r"Z\d[0-9A-Z]{6}_")
|
||||
MID_NAME = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
|
||||
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
|
||||
"sap_syntax_check", "sap_atc_run")
|
||||
GROUPS = [("official baseline Qwen (MacBook, 11 tasks)", "runs/stage1/baseline_qwen/*"),
|
||||
("Devstral baseline (dropped)", "runs/archive/devstral/baseline_devstral/*"),
|
||||
("older Qwen baselines (archive)", "runs/archive/qwen38/*/*"),
|
||||
("smoke", "runs/stage1/smoke/*"),
|
||||
("eval empirical filter (DeepSeek on eval candidates)", "runs/emp/*"),
|
||||
("eval generation validation (oracle, null, mutants)", "runs/gen/*"),
|
||||
("early pilot runs", "runs/2*_*"),
|
||||
("training trajectories (DeepSeek)", "runs/traj/*_G*"),
|
||||
("series A (local Qwen)", "runs/local_qwen/runs/*")]
|
||||
|
||||
|
||||
def own_prefix(d):
|
||||
m = re.match(r"^(\d+)_([GT]\d+)_", os.path.basename(d))
|
||||
return prefix_for(int(m.group(1)), m.group(2)).upper() if m else None
|
||||
|
||||
|
||||
def scan(d):
|
||||
own = own_prefix(d)
|
||||
p = os.path.join(d, "trajectory.jsonl")
|
||||
if not own or not os.path.exists(p):
|
||||
return None
|
||||
seen, reads, tools = set(), 0, Counter()
|
||||
for l in open(p):
|
||||
try:
|
||||
e = json.loads(l)
|
||||
except ValueError:
|
||||
continue
|
||||
if "tool" not in e:
|
||||
continue
|
||||
text = (e.get("result") or "").upper()
|
||||
found = {x for x in START.findall(text) if x != own and x[1].isdigit()}
|
||||
found |= {m.group(1).upper() for m in re.finditer(r"\b[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", text) if m.group(1).upper() != own}
|
||||
if found:
|
||||
seen |= found
|
||||
tools[e["tool"]] += 1
|
||||
name = str((e.get("args") or {}).get("objectName", "")).upper()
|
||||
m = START.match(name) or MID_NAME.match(name)
|
||||
pre = (m.group(1) if m and m.re is MID_NAME else (m.group(0) if m else "")).upper()
|
||||
if e["tool"] in READ and pre and pre != own:
|
||||
reads += 1
|
||||
return {"run": os.path.basename(d), "foreign_prefixes_seen": len(seen), "foreign_reads": reads, "tools": dict(tools)}
|
||||
|
||||
|
||||
def main():
|
||||
out = {}
|
||||
for label, pat in GROUPS:
|
||||
dirs = [d for d in sorted(glob.glob(os.path.join(ROOT, pat))) if os.path.isdir(d)]
|
||||
rows = [r for r in (scan(d) for d in dirs) if r]
|
||||
seen = [r for r in rows if r["foreign_prefixes_seen"]]
|
||||
reads = [r for r in rows if r["foreign_reads"]]
|
||||
out[label] = {"runs": len(rows), "runs_with_foreign_names": len(seen), "runs_with_foreign_reads": len(reads),
|
||||
"tools": dict(sum((Counter(r["tools"]) for r in seen), Counter())),
|
||||
"affected": [r["run"] for r in seen][:40], "read_runs": [r["run"] for r in reads]}
|
||||
print(f"{label}: {len(rows)} runs | foreign names in tool results: {len(seen)} | foreign reads: {len(reads)} | tools {out[label]['tools']}")
|
||||
os.makedirs(os.path.join(ROOT, "runs", "analysis"), exist_ok=True)
|
||||
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "foreign_scan.json"), "w"), indent=1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -5,10 +5,10 @@
|
||||
"""Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go.
|
||||
|
||||
Memory test first (a few dollars, finds the largest sequence length that fits, no data needed):
|
||||
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000
|
||||
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000,64000
|
||||
Real run (after the memory test and Kral's decision on the ratio):
|
||||
hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\
|
||||
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s1-epochs 1 --s2-epochs 3 --out erhankeseli/abap-mixed-adapter
|
||||
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s2-epochs 3 --s2-loss-share 0.6 --out erhankeseli/abap-mixed-adapter
|
||||
|
||||
Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns;
|
||||
none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut.
|
||||
@@ -35,6 +35,18 @@ def tokenize_masked(tok, text, spans, max_len):
|
||||
return ids, labels
|
||||
|
||||
|
||||
def s1_epochs_for(share, s2_loss_tokens, s2_epochs, s1_tokens):
|
||||
"""Epochs of stage 1 so that stage 2 carries `share` of all loss tokens (Kral + Opus 2026-10-06: share = 0.6)."""
|
||||
s2_total = s2_loss_tokens * s2_epochs
|
||||
return (1 - share) / share * s2_total / max(s1_tokens, 1)
|
||||
|
||||
|
||||
def copies(weight, rnd):
|
||||
"""Weight 1 = one copy per epoch, 2 = two, 0.5 = a copy in half of the epochs (a down-weight, not a drop)."""
|
||||
w = max(float(weight), 0.0)
|
||||
return int(w) + (1 if rnd.random() < w - int(w) else 0)
|
||||
|
||||
|
||||
def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
|
||||
rnd = random.Random(seed)
|
||||
rows = []
|
||||
@@ -45,7 +57,7 @@ def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
|
||||
rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))]
|
||||
for ep in range(int(s2_epochs)):
|
||||
for r in stage2:
|
||||
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * max(1, round(float(r.get("weight", 1.0))))
|
||||
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * copies(r.get("weight", 1.0), rnd)
|
||||
rnd.shuffle(rows)
|
||||
out, skipped = [], {"s1": 0, "s2": 0}
|
||||
for src, text, spans, _ in rows:
|
||||
@@ -65,8 +77,9 @@ def main():
|
||||
ap.add_argument("--s1-file", default="train.jsonl")
|
||||
ap.add_argument("--s2-file", default="stage2_train.jsonl")
|
||||
ap.add_argument("--s2-valid", default="stage2_valid.jsonl")
|
||||
ap.add_argument("--s1-epochs", type=float, default=1.0)
|
||||
ap.add_argument("--s1-epochs", type=float, default=None, help="epochs of stage 1; default: computed from --s2-loss-share")
|
||||
ap.add_argument("--s2-epochs", type=float, default=3.0)
|
||||
ap.add_argument("--s2-loss-share", type=float, default=0.6, help="share of the loss tokens that stage 2 carries (decision 2026-10-06: 0.6)")
|
||||
ap.add_argument("--max-seq", type=int, default=48000)
|
||||
ap.add_argument("--rank", type=int, default=16)
|
||||
ap.add_argument("--alpha", type=int, default=32)
|
||||
@@ -77,7 +90,7 @@ def main():
|
||||
ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter")
|
||||
ap.add_argument("--save-every", type=int, default=0)
|
||||
ap.add_argument("--memory-test", action="store_true")
|
||||
ap.add_argument("--sweep", default="16000,32000,48000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
|
||||
ap.add_argument("--sweep", default="16000,32000,48000,64000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
|
||||
ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length")
|
||||
a = ap.parse_args()
|
||||
|
||||
@@ -85,7 +98,8 @@ def main():
|
||||
from unsloth import FastLanguageModel
|
||||
from huggingface_hub import HfApi, hf_hub_download
|
||||
token = os.environ["HF_TOKEN"]
|
||||
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=a.max_seq, load_in_4bit=False, dtype=torch.bfloat16, token=token)
|
||||
seq_for_model = max([a.max_seq] + ([int(x) for x in a.sweep.split(",")] if a.memory_test else []))
|
||||
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=seq_for_model, load_in_4bit=False, dtype=torch.bfloat16, token=token)
|
||||
model = FastLanguageModel.get_peft_model(
|
||||
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
|
||||
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"],
|
||||
@@ -132,6 +146,11 @@ def main():
|
||||
p = hf_hub_download(repo, fname, repo_type="dataset", token=token)
|
||||
return [json.loads(l) for l in open(p)]
|
||||
s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file)
|
||||
if a.s1_epochs is None:
|
||||
s1_tok = sum(len(tok(r["text"], add_special_tokens=False)["input_ids"]) for r in s1)
|
||||
s2_loss = sum(r["assistant_tokens"] * float(r.get("weight", 1.0)) for r in s2)
|
||||
a.s1_epochs = s1_epochs_for(a.s2_loss_share, s2_loss, a.s2_epochs, s1_tok)
|
||||
print("STAGE1 EPOCHS", round(a.s1_epochs, 3), "(stage 1 tokens", s1_tok, ", stage 2 loss tokens per epoch", round(s2_loss), ", share", a.s2_loss_share, ")", flush=True)
|
||||
ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006)
|
||||
loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex)
|
||||
print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens",
|
||||
|
||||
@@ -10,17 +10,47 @@ def identity(row):
|
||||
return 1.0
|
||||
|
||||
|
||||
# Weights from the own-test mutation score (Kral + Opus 2026-10-06: down-weight, do not drop). Weight w means: w copies per epoch, a fraction is a
|
||||
# copy in that share of the epochs (0.5 = every second epoch on average). The acceptance filter itself is unchanged.
|
||||
OWN_TEST_WEIGHTS = {
|
||||
"reliable, score >= 0.75": 1.0, # the own tests pass on the correct reference and kill at least 3 of 4 mutants
|
||||
"reliable, score 0.5 to 0.75": 0.75,
|
||||
"reliable, score < 0.5": 0.5,
|
||||
"unreliable (own tests fail on the correct reference)": 0.5, # they may encode model specific behavior
|
||||
"no own tests": 0.5, # accepted at 85 points at most; the habit of writing tests is part of the behavior we want
|
||||
"no signal (PROG, no mutants, not scored)": 1.0,
|
||||
}
|
||||
|
||||
|
||||
def own_test_class(m):
|
||||
if m is None:
|
||||
return "no signal (PROG, no mutants, not scored)"
|
||||
st = m.get("status")
|
||||
if st == "no_own_tests":
|
||||
return "no own tests"
|
||||
if st != "scored":
|
||||
return "no signal (PROG, no mutants, not scored)"
|
||||
if not m.get("tests_pass_on_reference"):
|
||||
return "unreliable (own tests fail on the correct reference)"
|
||||
s = m.get("score")
|
||||
if s is None:
|
||||
return "no signal (PROG, no mutants, not scored)"
|
||||
return "reliable, score >= 0.75" if s >= 0.75 else "reliable, score 0.5 to 0.75" if s >= 0.5 else "reliable, score < 0.5"
|
||||
|
||||
|
||||
def own_test_weight(row):
|
||||
"""Item D (own-test mutation score, metadata only for now): reads runs/traj/<run>/own_test_mutation.json when it exists,
|
||||
stores it as extra data and does NOT drop or reweight (Kral + Opus 2026-10-06: do not change the acceptance yet)."""
|
||||
"""Item D: reads runs/traj/<run>/own_test_mutation.json, stores the score as extra data and sets the weight by OWN_TEST_WEIGHTS."""
|
||||
run = row["id"].split("_r")[-1]
|
||||
m = None
|
||||
for d in os.listdir(os.path.join(ROOT, "runs", "traj")):
|
||||
if d.startswith(run + "_"):
|
||||
p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json")
|
||||
if os.path.exists(p):
|
||||
m = json.load(open(p))
|
||||
return {"keep": True, "weight": 1.0, "extra": {"own_test_mutation": m.get("score"), "own_test_mutants": m.get("mutants")}}
|
||||
return {"keep": True, "weight": 1.0}
|
||||
break
|
||||
cls = own_test_class(m)
|
||||
return {"keep": True, "weight": OWN_TEST_WEIGHTS[cls],
|
||||
"extra": {"own_test_class": cls, "own_test_mutation": (m or {}).get("score"), "own_test_reliable": bool((m or {}).get("tests_pass_on_reference"))}}
|
||||
|
||||
|
||||
def repair_up(row):
|
||||
|
||||
71
train/mem_table.py
Normal file
71
train/mem_table.py
Normal file
@@ -0,0 +1,71 @@
|
||||
"""Writes docs/bf16-memory.md (estimates for the bf16 LoRA run of Qwen 3.8 27B by GPU and sequence length)."""
|
||||
import json
|
||||
import os
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
H, L, V, I = 5120, 64, 248320, 17408
|
||||
r = 16
|
||||
lora = 16 * r * ((H + 6144) + (H + 1024) + (H + 1024) + (6144 + H)) + 48 * r * ((H + 10240) + (H + 6144) + (6144 + H)) + L * r * ((H + I) * 2 + (I + H))
|
||||
W = 27.8e9 * 2 / 2**30
|
||||
lora_gb = lora * 10 / 2**30
|
||||
|
||||
|
||||
def est(seq, offload, chunked):
|
||||
ckpt = 0.0 if offload else seq * H * 2 * L / 2**30
|
||||
layer = seq * I * 2 * 4 / 2**30 + 3
|
||||
logits = (seq * V * 2 * 3 / 2**30) if not chunked else 3.0
|
||||
return W + lora_gb + ckpt + layer + logits, dict(weights=W, lora=lora_gb, ckpt=ckpt, layer=layer, logits=logits)
|
||||
|
||||
|
||||
gpus = [("a100-large", "1x A100 80 GB", 80), ("rtx-pro-6000", "1x RTX PRO 6000 96 GB", 96), ("h200", "1x H200 141 GB", 141), ("h200x2", "2x H200 282 GB (model parallel)", 282)]
|
||||
tab = "| seq | checkpoints | loss | estimated peak GB | " + " | ".join(g[1] for g in gpus) + " |\n|---|---|---|---|" + "---|" * len(gpus) + "\n"
|
||||
for seq in (16000, 32000, 48000, 64000):
|
||||
for off in (True, False):
|
||||
for ch in (True, False):
|
||||
t, _ = est(seq, off, ch)
|
||||
cells = ["fits" if t <= g[2] * 0.92 else ("tight" if t <= g[2] else "no") for g in gpus]
|
||||
tab += f"| {seq // 1000}k | {'CPU offload (Unsloth)' if off else 'on the GPU'} | {'chunked' if ch else 'full logits'} | {t:.0f} | " + " | ".join(cells) + " |\n"
|
||||
_, b48 = est(48000, True, True)
|
||||
_, b64 = est(64000, True, True)
|
||||
doc = f"""# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
|
||||
|
||||
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
|
||||
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
|
||||
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
|
||||
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
|
||||
|
||||
## Model (from config.json of Qwen3.8-27B)
|
||||
64 layers (48 linear attention, 16 full attention), hidden {H}, MLP {I}, vocab {V}, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **{lora / 1e6:.0f} M trainable parameters**.
|
||||
|
||||
## Memory by component (GB; estimate, not a measurement)
|
||||
| component | 48k | 64k | how |
|
||||
|---|---|---|---|
|
||||
| weights bf16 | {b48['weights']:.1f} | {b64['weights']:.1f} | 27.8 B x 2 bytes |
|
||||
| LoRA weights + grads + 8-bit Adam | {b48['lora']:.1f} | {b64['lora']:.1f} | {lora / 1e6:.0f} M x 10 bytes |
|
||||
| layer inputs for the backward pass | {48000 * H * 2 * L / 2**30:.1f} on the GPU, about 0 with the Unsloth CPU offload (then {48000 * H * 2 * L / 2**30:.0f} GB host RAM) | {64000 * H * 2 * L / 2**30:.1f} / about 0 (host RAM {64000 * H * 2 * L / 2**30:.0f} GB) | seq x 5120 x 2 bytes x 64 layers |
|
||||
| recompute peak of one layer | {b48['layer']:.1f} | {b64['layer']:.1f} | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
|
||||
| logits and loss | {b48['logits']:.0f} chunked, {48000 * V * 2 * 3 / 2**30:.0f} full logits | {b64['logits']:.0f} chunked, {64000 * V * 2 * 3 / 2**30:.0f} full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
|
||||
|
||||
## Fit by GPU (92 % of the card counted as usable)
|
||||
{tab}
|
||||
Reading:
|
||||
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
|
||||
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
|
||||
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
|
||||
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about {b48['weights'] + b48['lora'] + b48['layer'] + b48['logits']:.0f} GB at 48k, {b64['weights'] + b64['lora'] + b64['layer'] + b64['logits']:.0f} GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
|
||||
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
|
||||
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
|
||||
|
||||
## Order when the test is run
|
||||
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
|
||||
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
|
||||
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
|
||||
|
||||
## Time and cost (estimate from the nf4 run: {6750 / 30.5:.0f} tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
|
||||
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
|
||||
A100 about {3.3e6 / 300 / 3600:.1f} h (about {3.3e6 / 300 / 3600 * 2.5:.0f} USD), H200 about {3.3e6 / 650 / 3600:.1f} h (about {3.3e6 / 650 / 3600 * 5:.0f} USD).
|
||||
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
|
||||
H200 about {10e6 / 650 / 3600:.1f} h (about {10e6 / 650 / 3600 * 5:.0f} USD), A100 about {10e6 / 300 / 3600:.0f} h (about {10e6 / 300 / 3600 * 2.5:.0f} USD). The real number comes from the memory test (`step_seconds`).
|
||||
"""
|
||||
open(os.path.join(ROOT, "docs", "bf16-memory.md"), "w").write(doc)
|
||||
print("written")
|
||||
75
train/requeue_foreign.py
Normal file
75
train/requeue_foreign.py
Normal file
@@ -0,0 +1,75 @@
|
||||
"""Trajectories in which the model read another run's leftover object go back to the pending pool (Kral + Opus 2026-10-06).
|
||||
|
||||
python3 train/requeue_foreign.py dry run
|
||||
python3 train/requeue_foreign.py --apply moves their rows from runs/traj/summary.jsonl to runs/traj/summary_excluded.jsonl
|
||||
(backup: summary.jsonl.bak-<time>); the task then counts as not run and is run again
|
||||
after the reset in the normal order (kind deficit). The run folders stay.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sys
|
||||
import time
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
sys.path.insert(0, ROOT)
|
||||
sys.path.insert(0, os.path.join(ROOT, "train"))
|
||||
import accept as acc # noqa: E402
|
||||
from harness import mix # noqa: E402
|
||||
|
||||
START = __import__("re").compile(r"^(Z\d[0-9A-Z]{6}_)", __import__("re").I)
|
||||
MID = __import__("re").compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", __import__("re").I)
|
||||
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
|
||||
"sap_syntax_check", "sap_atc_run")
|
||||
|
||||
|
||||
def foreign_reads(rec):
|
||||
own = rec["prefix"].upper()
|
||||
n = 0
|
||||
for m in rec["messages"]:
|
||||
if m["role"] != "assistant":
|
||||
continue
|
||||
for c in m.get("tool_calls") or []:
|
||||
if c["function"]["name"] not in READ:
|
||||
continue
|
||||
a = c["function"].get("arguments") or "{}"
|
||||
try:
|
||||
a = json.loads(a) if isinstance(a, str) else a
|
||||
except ValueError:
|
||||
a = {}
|
||||
name = str(a.get("objectName", "")).upper()
|
||||
mm = START.match(name) or MID.match(name)
|
||||
if mm and mm.group(1).upper() != own:
|
||||
n += 1
|
||||
return n
|
||||
|
||||
|
||||
def main():
|
||||
apply = "--apply" in sys.argv
|
||||
path = os.path.join(ROOT, "runs", "traj", "summary.jsonl")
|
||||
rows = [json.loads(l) for l in open(path)]
|
||||
keep, out = [], []
|
||||
for r in rows:
|
||||
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
|
||||
if os.path.exists(p):
|
||||
rec = json.load(open(p))
|
||||
if acc.judge(rec, r, 80)[0]:
|
||||
n = foreign_reads(rec)
|
||||
if n:
|
||||
out.append(dict(r, excluded_reason="read of another run's object", foreign_reads=n, kind=mix.kind_of_task_dir(r["task"]),
|
||||
excluded_at=time.strftime("%F %T")))
|
||||
continue
|
||||
keep.append(r)
|
||||
print(len(out), "accepted trajectories with a foreign read:", [(o["task"], o["attempt"], o["kind"], o["foreign_reads"]) for o in out])
|
||||
if not apply:
|
||||
return
|
||||
shutil.copy(path, path + ".bak-" + time.strftime("%Y%m%d-%H%M%S"))
|
||||
with open(os.path.join(ROOT, "runs", "traj", "summary_excluded.jsonl"), "a") as f:
|
||||
for o in out:
|
||||
f.write(json.dumps(o) + "\n")
|
||||
open(path, "w").write("".join(json.dumps(r) + "\n" for r in keep))
|
||||
print("moved; summary.jsonl now", len(keep), "rows")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user