Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 09:04:09 +02:00
parent a465c33e3f
commit 8ef85c2713
10 changed files with 402 additions and 41 deletions

View File

@@ -43,7 +43,7 @@ Loss mask checked on every train sample: no span contains a tool result, the sys
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides)
## Stage 1 : stage 2 ratio (DECIDED 2026-10-06: the 60 % rule)
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
@@ -61,7 +61,28 @@ Reasons: (1) stage 2 is the behavior we want (repair after the first error: the
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
**Decision (Kral + Opus 2026-10-06): this rule.** `train/hf_train_bf16.py --s2-loss-share 0.6` (default) computes the stage 1 epochs from the weighted stage 2 loss tokens.
## Own-test weights (decided 2026-10-06: down-weight, do not drop; `train/hooks_example.py`, `own_test_weight`)
A weight w is the number of copies per epoch (0.5 = a copy in every second epoch on average). The acceptance filter is unchanged.
| class of the trajectory | weight | reason |
|---|---|---|
| own tests pass on the correct reference and kill 3 of 4 mutants or more (score >= 0.75) | 1.0 | strong tests, the behavior we want |
| reliable, score 0.5 to 0.75 | 0.75 | weaker tests |
| reliable, score under 0.5 | 0.5 | tests that miss most faults |
| own tests fail on the correct reference (unreliable) | 0.5 | may encode model specific behavior |
| no own tests (accepted at 85 points at most) | 0.5 | writing tests is part of the behavior to teach |
| no signal (PROG, no mutants, not scored) | 1.0 | neither good nor bad |
In today's 36 train samples: 15 reliable at 1.0, 4 without signal at 1.0, **15 without own tests at 0.5, 2 unreliable at 0.5**: the expected stage 2 loss tokens per epoch fall from 0.32 M to 0.24 M, which gives about 0.2 epochs of stage 1.
Check at the real build: the CLAS cap picks repair trajectories first, and many of them have no own tests; if the share of down-weighted samples stays near half, the weights need a second look.
## Other decisions of 2026-10-06
- **CLAS cap 35 %** stays a build setting (`--clas-cap`); review at the real build (with so little non-CLAS data it throws good CLAS samples into the reserve).
- **Length limit:** the memory test decides (`docs/bf16-memory.md`: sweep 16k, 32k, 48k, 64k); the builder keeps 48k until then.
- **Foreign names:** list results scrubbed; trajectories with a foreign read are dropped and their tasks go back to the pending pool (`docs/foreign-objects-report.md`).
- **Memory test:** waits until the data is near the size of the first SFT run.
## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).

79
train/foreign_scan.py Normal file
View File

@@ -0,0 +1,79 @@
"""Read-only scan: which runs saw other runs' objects in tool results, or read them (teardown/proxy bug of 2026-10-06).
python3 train/foreign_scan.py -> prints a summary and writes runs/analysis/foreign_scan.json
"""
import glob
import json
import os
import re
import sys
from collections import Counter
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
from harness.task import prefix_for # noqa: E402
START = re.compile(r"Z\d[0-9A-Z]{6}_")
MID_NAME = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
GROUPS = [("official baseline Qwen (MacBook, 11 tasks)", "runs/stage1/baseline_qwen/*"),
("Devstral baseline (dropped)", "runs/archive/devstral/baseline_devstral/*"),
("older Qwen baselines (archive)", "runs/archive/qwen38/*/*"),
("smoke", "runs/stage1/smoke/*"),
("eval empirical filter (DeepSeek on eval candidates)", "runs/emp/*"),
("eval generation validation (oracle, null, mutants)", "runs/gen/*"),
("early pilot runs", "runs/2*_*"),
("training trajectories (DeepSeek)", "runs/traj/*_G*"),
("series A (local Qwen)", "runs/local_qwen/runs/*")]
def own_prefix(d):
m = re.match(r"^(\d+)_([GT]\d+)_", os.path.basename(d))
return prefix_for(int(m.group(1)), m.group(2)).upper() if m else None
def scan(d):
own = own_prefix(d)
p = os.path.join(d, "trajectory.jsonl")
if not own or not os.path.exists(p):
return None
seen, reads, tools = set(), 0, Counter()
for l in open(p):
try:
e = json.loads(l)
except ValueError:
continue
if "tool" not in e:
continue
text = (e.get("result") or "").upper()
found = {x for x in START.findall(text) if x != own and x[1].isdigit()}
found |= {m.group(1).upper() for m in re.finditer(r"\b[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", text) if m.group(1).upper() != own}
if found:
seen |= found
tools[e["tool"]] += 1
name = str((e.get("args") or {}).get("objectName", "")).upper()
m = START.match(name) or MID_NAME.match(name)
pre = (m.group(1) if m and m.re is MID_NAME else (m.group(0) if m else "")).upper()
if e["tool"] in READ and pre and pre != own:
reads += 1
return {"run": os.path.basename(d), "foreign_prefixes_seen": len(seen), "foreign_reads": reads, "tools": dict(tools)}
def main():
out = {}
for label, pat in GROUPS:
dirs = [d for d in sorted(glob.glob(os.path.join(ROOT, pat))) if os.path.isdir(d)]
rows = [r for r in (scan(d) for d in dirs) if r]
seen = [r for r in rows if r["foreign_prefixes_seen"]]
reads = [r for r in rows if r["foreign_reads"]]
out[label] = {"runs": len(rows), "runs_with_foreign_names": len(seen), "runs_with_foreign_reads": len(reads),
"tools": dict(sum((Counter(r["tools"]) for r in seen), Counter())),
"affected": [r["run"] for r in seen][:40], "read_runs": [r["run"] for r in reads]}
print(f"{label}: {len(rows)} runs | foreign names in tool results: {len(seen)} | foreign reads: {len(reads)} | tools {out[label]['tools']}")
os.makedirs(os.path.join(ROOT, "runs", "analysis"), exist_ok=True)
json.dump(out, open(os.path.join(ROOT, "runs", "analysis", "foreign_scan.json"), "w"), indent=1)
if __name__ == "__main__":
main()

View File

@@ -5,10 +5,10 @@
"""Stage 1 + stage 2 mixed bf16 LoRA run for Qwen 3.8 27B on Hugging Face Jobs (Opus item E, 2026-10-06). NOT started: no job without Kral's go.
Memory test first (a few dollars, finds the largest sequence length that fits, no data needed):
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000
hf jobs uv run --flavor h200 --timeout 40m --secrets HF_TOKEN train/hf_train_bf16.py -- --memory-test --sweep 16000,32000,48000,64000
Real run (after the memory test and Kral's decision on the ratio):
hf jobs uv run --flavor <flavor> --timeout 10h --secrets HF_TOKEN train/hf_train_bf16.py -- \\
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s1-epochs 1 --s2-epochs 3 --out erhankeseli/abap-mixed-adapter
--stage1 erhankeseli/abap-stage1-data --stage2 erhankeseli/abap-stage2-data --s2-epochs 3 --s2-loss-share 0.6 --out erhankeseli/abap-mixed-adapter
Data: stage 1 rows have `text` (loss on every token); stage 2 rows have `text` and `assistant_spans` (loss only inside the spans: assistant turns;
none on the system turn with the tool schemas, the user turn or the tool results). A sample longer than --max-seq is skipped, never cut.
@@ -35,6 +35,18 @@ def tokenize_masked(tok, text, spans, max_len):
return ids, labels
def s1_epochs_for(share, s2_loss_tokens, s2_epochs, s1_tokens):
"""Epochs of stage 1 so that stage 2 carries `share` of all loss tokens (Kral + Opus 2026-10-06: share = 0.6)."""
s2_total = s2_loss_tokens * s2_epochs
return (1 - share) / share * s2_total / max(s1_tokens, 1)
def copies(weight, rnd):
"""Weight 1 = one copy per epoch, 2 = two, 0.5 = a copy in half of the epochs (a down-weight, not a drop)."""
w = max(float(weight), 0.0)
return int(w) + (1 if rnd.random() < w - int(w) else 0)
def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
rnd = random.Random(seed)
rows = []
@@ -45,7 +57,7 @@ def build_examples(tok, stage1, stage2, s1_epochs, s2_epochs, max_len, seed):
rows += [("s1", r["text"], None, 1.0) for r in rnd.sample(stage1, int(len(stage1) * frac))]
for ep in range(int(s2_epochs)):
for r in stage2:
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * max(1, round(float(r.get("weight", 1.0))))
rows += [("s2", r["text"], r["assistant_spans"], 1.0)] * copies(r.get("weight", 1.0), rnd)
rnd.shuffle(rows)
out, skipped = [], {"s1": 0, "s2": 0}
for src, text, spans, _ in rows:
@@ -65,8 +77,9 @@ def main():
ap.add_argument("--s1-file", default="train.jsonl")
ap.add_argument("--s2-file", default="stage2_train.jsonl")
ap.add_argument("--s2-valid", default="stage2_valid.jsonl")
ap.add_argument("--s1-epochs", type=float, default=1.0)
ap.add_argument("--s1-epochs", type=float, default=None, help="epochs of stage 1; default: computed from --s2-loss-share")
ap.add_argument("--s2-epochs", type=float, default=3.0)
ap.add_argument("--s2-loss-share", type=float, default=0.6, help="share of the loss tokens that stage 2 carries (decision 2026-10-06: 0.6)")
ap.add_argument("--max-seq", type=int, default=48000)
ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--alpha", type=int, default=32)
@@ -77,7 +90,7 @@ def main():
ap.add_argument("--out", default="erhankeseli/abap-mixed-adapter")
ap.add_argument("--save-every", type=int, default=0)
ap.add_argument("--memory-test", action="store_true")
ap.add_argument("--sweep", default="16000,32000,48000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
ap.add_argument("--sweep", default="16000,32000,48000,64000", help="memory test: sequence lengths, tried in this order, stops at the first OOM")
ap.add_argument("--steps", type=int, default=3, help="memory test: optimizer steps per length")
a = ap.parse_args()
@@ -85,7 +98,8 @@ def main():
from unsloth import FastLanguageModel
from huggingface_hub import HfApi, hf_hub_download
token = os.environ["HF_TOKEN"]
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=a.max_seq, load_in_4bit=False, dtype=torch.bfloat16, token=token)
seq_for_model = max([a.max_seq] + ([int(x) for x in a.sweep.split(",")] if a.memory_test else []))
model, tok = FastLanguageModel.from_pretrained(a.model, max_seq_length=seq_for_model, load_in_4bit=False, dtype=torch.bfloat16, token=token)
model = FastLanguageModel.get_peft_model(
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", "in_proj_qkv", "in_proj_z", "out_proj"],
@@ -132,6 +146,11 @@ def main():
p = hf_hub_download(repo, fname, repo_type="dataset", token=token)
return [json.loads(l) for l in open(p)]
s1, s2 = rows(a.stage1, a.s1_file), rows(a.stage2, a.s2_file)
if a.s1_epochs is None:
s1_tok = sum(len(tok(r["text"], add_special_tokens=False)["input_ids"]) for r in s1)
s2_loss = sum(r["assistant_tokens"] * float(r.get("weight", 1.0)) for r in s2)
a.s1_epochs = s1_epochs_for(a.s2_loss_share, s2_loss, a.s2_epochs, s1_tok)
print("STAGE1 EPOCHS", round(a.s1_epochs, 3), "(stage 1 tokens", s1_tok, ", stage 2 loss tokens per epoch", round(s2_loss), ", share", a.s2_loss_share, ")", flush=True)
ex, skipped = build_examples(tok, s1, s2, a.s1_epochs, a.s2_epochs, a.max_seq, 20261006)
loss_tok = sum(sum(1 for x in e["labels"] if x != -100) for e in ex)
print("EXAMPLES", len(ex), "skipped", skipped, "loss tokens", loss_tok, "stage 2 share of loss tokens",

View File

@@ -10,17 +10,47 @@ def identity(row):
return 1.0
# Weights from the own-test mutation score (Kral + Opus 2026-10-06: down-weight, do not drop). Weight w means: w copies per epoch, a fraction is a
# copy in that share of the epochs (0.5 = every second epoch on average). The acceptance filter itself is unchanged.
OWN_TEST_WEIGHTS = {
"reliable, score >= 0.75": 1.0, # the own tests pass on the correct reference and kill at least 3 of 4 mutants
"reliable, score 0.5 to 0.75": 0.75,
"reliable, score < 0.5": 0.5,
"unreliable (own tests fail on the correct reference)": 0.5, # they may encode model specific behavior
"no own tests": 0.5, # accepted at 85 points at most; the habit of writing tests is part of the behavior we want
"no signal (PROG, no mutants, not scored)": 1.0,
}
def own_test_class(m):
if m is None:
return "no signal (PROG, no mutants, not scored)"
st = m.get("status")
if st == "no_own_tests":
return "no own tests"
if st != "scored":
return "no signal (PROG, no mutants, not scored)"
if not m.get("tests_pass_on_reference"):
return "unreliable (own tests fail on the correct reference)"
s = m.get("score")
if s is None:
return "no signal (PROG, no mutants, not scored)"
return "reliable, score >= 0.75" if s >= 0.75 else "reliable, score 0.5 to 0.75" if s >= 0.5 else "reliable, score < 0.5"
def own_test_weight(row):
"""Item D (own-test mutation score, metadata only for now): reads runs/traj/<run>/own_test_mutation.json when it exists,
stores it as extra data and does NOT drop or reweight (Kral + Opus 2026-10-06: do not change the acceptance yet)."""
"""Item D: reads runs/traj/<run>/own_test_mutation.json, stores the score as extra data and sets the weight by OWN_TEST_WEIGHTS."""
run = row["id"].split("_r")[-1]
m = None
for d in os.listdir(os.path.join(ROOT, "runs", "traj")):
if d.startswith(run + "_"):
p = os.path.join(ROOT, "runs", "traj", d, "own_test_mutation.json")
if os.path.exists(p):
m = json.load(open(p))
return {"keep": True, "weight": 1.0, "extra": {"own_test_mutation": m.get("score"), "own_test_mutants": m.get("mutants")}}
return {"keep": True, "weight": 1.0}
break
cls = own_test_class(m)
return {"keep": True, "weight": OWN_TEST_WEIGHTS[cls],
"extra": {"own_test_class": cls, "own_test_mutation": (m or {}).get("score"), "own_test_reliable": bool((m or {}).get("tests_pass_on_reference"))}}
def repair_up(row):

71
train/mem_table.py Normal file
View File

@@ -0,0 +1,71 @@
"""Writes docs/bf16-memory.md (estimates for the bf16 LoRA run of Qwen 3.8 27B by GPU and sequence length)."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
H, L, V, I = 5120, 64, 248320, 17408
r = 16
lora = 16 * r * ((H + 6144) + (H + 1024) + (H + 1024) + (6144 + H)) + 48 * r * ((H + 10240) + (H + 6144) + (6144 + H)) + L * r * ((H + I) * 2 + (I + H))
W = 27.8e9 * 2 / 2**30
lora_gb = lora * 10 / 2**30
def est(seq, offload, chunked):
ckpt = 0.0 if offload else seq * H * 2 * L / 2**30
layer = seq * I * 2 * 4 / 2**30 + 3
logits = (seq * V * 2 * 3 / 2**30) if not chunked else 3.0
return W + lora_gb + ckpt + layer + logits, dict(weights=W, lora=lora_gb, ckpt=ckpt, layer=layer, logits=logits)
gpus = [("a100-large", "1x A100 80 GB", 80), ("rtx-pro-6000", "1x RTX PRO 6000 96 GB", 96), ("h200", "1x H200 141 GB", 141), ("h200x2", "2x H200 282 GB (model parallel)", 282)]
tab = "| seq | checkpoints | loss | estimated peak GB | " + " | ".join(g[1] for g in gpus) + " |\n|---|---|---|---|" + "---|" * len(gpus) + "\n"
for seq in (16000, 32000, 48000, 64000):
for off in (True, False):
for ch in (True, False):
t, _ = est(seq, off, ch)
cells = ["fits" if t <= g[2] * 0.92 else ("tight" if t <= g[2] else "no") for g in gpus]
tab += f"| {seq // 1000}k | {'CPU offload (Unsloth)' if off else 'on the GPU'} | {'chunked' if ch else 'full logits'} | {t:.0f} | " + " | ".join(cells) + " |\n"
_, b48 = est(48000, True, True)
_, b64 = est(64000, True, True)
doc = f"""# bf16 memory test and GPU choice (2026-10-06, estimates, nothing was run)
**No GPU job is started without Kral's go. Decision 2026-10-06: wait. Run the memory test when the training data is near the size of the first SFT run** (the real sample lengths and the
count then decide the flavor; today there are 36 train samples). Script: `train/hf_train_bf16.py` (`--memory-test --sweep 16000,32000,48000,64000`: no data needed, 3 optimizer steps per length,
stops at the first OOM, uploads the result as `memtest_*.json` to the output repo). **The length limit (48k or 64k) is decided by this test** (decision 3): the builder keeps `--max-tokens 48000`
until then; 64k would bring back the long CDS trajectories (today 5 of 19 are over 48k).
## Model (from config.json of Qwen3.8-27B)
64 layers (48 linear attention, 16 full attention), hidden {H}, MLP {I}, vocab {V}, 27.8 B parameters. LoRA rank 16 on q/k/v/o, gate/up/down, in_proj_qkv, in_proj_z, out_proj: **{lora / 1e6:.0f} M trainable parameters**.
## Memory by component (GB; estimate, not a measurement)
| component | 48k | 64k | how |
|---|---|---|---|
| weights bf16 | {b48['weights']:.1f} | {b64['weights']:.1f} | 27.8 B x 2 bytes |
| LoRA weights + grads + 8-bit Adam | {b48['lora']:.1f} | {b64['lora']:.1f} | {lora / 1e6:.0f} M x 10 bytes |
| layer inputs for the backward pass | {48000 * H * 2 * L / 2**30:.1f} on the GPU, about 0 with the Unsloth CPU offload (then {48000 * H * 2 * L / 2**30:.0f} GB host RAM) | {64000 * H * 2 * L / 2**30:.1f} / about 0 (host RAM {64000 * H * 2 * L / 2**30:.0f} GB) | seq x 5120 x 2 bytes x 64 layers |
| recompute peak of one layer | {b48['layer']:.1f} | {b64['layer']:.1f} | MLP tensors seq x 17408 x 2 bytes x 4 plus 3 GB for attention (assumption) |
| logits and loss | {b48['logits']:.0f} chunked, {48000 * V * 2 * 3 / 2**30:.0f} full logits | {b64['logits']:.0f} chunked, {64000 * V * 2 * 3 / 2**30:.0f} full logits | full logits: seq x 248320 x (bf16 + fp32 upcast + grad) |
## Fit by GPU (92 % of the card counted as usable)
{tab}
Reading:
- **The loss is the main risk, not the weights.** With full logits only the H200 (141 GB, CPU offload) fits 32k and 48k, and not 64k. The run needs a fused or chunked cross entropy
(Unsloth has one for the architectures it patches; whether it covers `qwen3_5` is unknown). The memory test shows it at once: if 16000 already fails on an H200, the loss is the cause.
Fallback (not built): hidden states, then the loss over the 32 % labeled positions only, in chunks of 4k.
- **A100 80 GB (2.50 USD/hour)** only with the CPU offload and a chunked loss (about {b48['weights'] + b48['lora'] + b48['layer'] + b48['logits']:.0f} GB at 48k, {b64['weights'] + b64['lora'] + b64['layer'] + b64['logits']:.0f} GB at 64k: no margin). **H200 141 GB (5 USD/hour)** fits with margin;
**RTX PRO 6000 96 GB (2.75 USD/hour)** fits with the offload and the chunked loss.
- Multi-GPU (`a100x4`, `h200x2`): only with model parallelism, one card works at a time; not before the single card test.
## Order when the test is run
1. Memory test on **h200** (about 40 minutes with 64k, about 3.5 USD): sweep 16000,32000,48000,64000, offload on. The limit is the largest length that fits with margin.
2. If the loss is the problem: build the chunked loss (about 2 hours), repeat.
3. Real run with `--s2-epochs 3 --s2-loss-share 0.6` (stage 1 epochs are computed: stage 2 carries 60 % of the loss tokens; weights of the own-test class are applied).
## Time and cost (estimate from the nf4 run: {6750 / 30.5:.0f} tokens/s on an A100 at 6.75k tokens per document; bf16 faster, an H200 about 2 to 2.5 times an A100)
Today's data: stage 2 is 36 samples (0.94 M tokens per epoch, 3 epochs = 2.8 M tokens) and, by the 60 % rule, about 0.2 epochs of stage 1 (0.5 M tokens): 3.3 M tokens in total.
A100 about {3.3e6 / 300 / 3600:.1f} h (about {3.3e6 / 300 / 3600 * 2.5:.0f} USD), H200 about {3.3e6 / 650 / 3600:.1f} h (about {3.3e6 / 650 / 3600 * 5:.0f} USD).
With three times the stage 2 data (the size that the restart plan aims at): 2.8 M tokens per epoch, 8.5 M in 3 epochs, plus about 0.6 epochs of stage 1 (1.5 M): 10 M tokens,
H200 about {10e6 / 650 / 3600:.1f} h (about {10e6 / 650 / 3600 * 5:.0f} USD), A100 about {10e6 / 300 / 3600:.0f} h (about {10e6 / 300 / 3600 * 2.5:.0f} USD). The real number comes from the memory test (`step_seconds`).
"""
open(os.path.join(ROOT, "docs", "bf16-memory.md"), "w").write(doc)
print("written")

75
train/requeue_foreign.py Normal file
View File

@@ -0,0 +1,75 @@
"""Trajectories in which the model read another run's leftover object go back to the pending pool (Kral + Opus 2026-10-06).
python3 train/requeue_foreign.py dry run
python3 train/requeue_foreign.py --apply moves their rows from runs/traj/summary.jsonl to runs/traj/summary_excluded.jsonl
(backup: summary.jsonl.bak-<time>); the task then counts as not run and is run again
after the reset in the normal order (kind deficit). The run folders stay.
"""
import json
import os
import shutil
import sys
import time
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, ROOT)
sys.path.insert(0, os.path.join(ROOT, "train"))
import accept as acc # noqa: E402
from harness import mix # noqa: E402
START = __import__("re").compile(r"^(Z\d[0-9A-Z]{6}_)", __import__("re").I)
MID = __import__("re").compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", __import__("re").I)
READ = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
def foreign_reads(rec):
own = rec["prefix"].upper()
n = 0
for m in rec["messages"]:
if m["role"] != "assistant":
continue
for c in m.get("tool_calls") or []:
if c["function"]["name"] not in READ:
continue
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
name = str(a.get("objectName", "")).upper()
mm = START.match(name) or MID.match(name)
if mm and mm.group(1).upper() != own:
n += 1
return n
def main():
apply = "--apply" in sys.argv
path = os.path.join(ROOT, "runs", "traj", "summary.jsonl")
rows = [json.loads(l) for l in open(path)]
keep, out = [], []
for r in rows:
p = os.path.join(ROOT, "runs", "traj", r.get("run_dir") or "-", "record.json")
if os.path.exists(p):
rec = json.load(open(p))
if acc.judge(rec, r, 80)[0]:
n = foreign_reads(rec)
if n:
out.append(dict(r, excluded_reason="read of another run's object", foreign_reads=n, kind=mix.kind_of_task_dir(r["task"]),
excluded_at=time.strftime("%F %T")))
continue
keep.append(r)
print(len(out), "accepted trajectories with a foreign read:", [(o["task"], o["attempt"], o["kind"], o["foreign_reads"]) for o in out])
if not apply:
return
shutil.copy(path, path + ".bak-" + time.strftime("%Y%m%d-%H%M%S"))
with open(os.path.join(ROOT, "runs", "traj", "summary_excluded.jsonl"), "a") as f:
for o in out:
f.write(json.dumps(o) + "\n")
open(path, "w").write("".join(json.dumps(r) + "\n" for r in keep))
print("moved; summary.jsonl now", len(keep), "rows")
if __name__ == "__main__":
main()