stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
70
train/build_doc.py
Normal file
70
train/build_doc.py
Normal file
@@ -0,0 +1,70 @@
|
||||
"""Writes docs/stage2-build.md from runs/stage2_data/build_report.json (run after train/build_stage2.py)."""
|
||||
import json
|
||||
import os
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
rep = json.load(open(os.path.join(ROOT, "runs", "stage2_data", "build_report.json")))
|
||||
tr, va, rs = rep["train"], rep["valid"], rep["reserve"]
|
||||
s1 = rep["stage1_stats"]["train"]["tokens"]
|
||||
s1d = rep["stage1_stats"]["train"]["docs"]
|
||||
L2 = tr["loss_tokens"]
|
||||
|
||||
|
||||
def row(e1, e2, L=L2):
|
||||
a, b = s1 * e1, L * e2
|
||||
return f"| {e1} | {e2} | {a / 1e6:.2f} M | {b / 1e6:.2f} M | {100 * b / (a + b):.0f} % | {(a + b) / 1e6:.1f} M |"
|
||||
|
||||
|
||||
drop = rep["dropped"]
|
||||
dropped_lines = "\n".join(f"- {k}: {v}" for k, v in drop.items())
|
||||
kinds = "\n".join(f"| {k} | {v['samples']} | {v['families']} | {v['short']} |" for k, v in rep["kinds"].items())
|
||||
doc = f"""# Stage 2 training set, build of {rep['built']}
|
||||
|
||||
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
||||
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||
|
||||
## Result
|
||||
{rep['steps']['accepted_trajectories']} accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> {rep['steps']['after_eval_overlap']} after the eval overlap check
|
||||
-> {rep['steps']['after_token_limit']} after the 48k limit (none cut) -> CLAS cap 35 %: {rep['steps']['clas_cap']['clas_kept']} CLAS kept, **{rep['steps']['clas_cap']['clas_reserve']} CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||
-> **{tr['samples']} train + {va['samples']} valid samples** ({tr['tokens'] / 1e6:.2f} M + {va['tokens'] / 1e6:.2f} M tokens, {tr['loss_tokens'] / 1e6:.2f} M loss tokens in train, p50 {tr['p50']}, p95 {tr['p95']}, max {tr['max']}, repair share {tr['repair_share']}).
|
||||
Dropped:
|
||||
{dropped_lines}
|
||||
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
||||
|
||||
## Not good enough yet (kinds under the minimum of {rep['settings']['min_per_kind']})
|
||||
| kind | samples | families | short |
|
||||
|---|---|---|---|
|
||||
{kinds}
|
||||
|
||||
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
|
||||
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
|
||||
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
|
||||
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
|
||||
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||
|
||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
|
||||
|
||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||
|---|---|---|---|---|---|
|
||||
{row(2, 3)}
|
||||
{row(1, 3)}
|
||||
{row(0.5, 3)}
|
||||
{row(0.33, 3)}
|
||||
{row(1, 3, 3 * L2)}
|
||||
|
||||
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
|
||||
|
||||
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
|
||||
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
|
||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||
|
||||
## Also built
|
||||
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||
"""
|
||||
open(os.path.join(ROOT, "docs", "stage2-build.md"), "w").write(doc)
|
||||
print("written", len(doc))
|
||||
@@ -14,6 +14,7 @@ nothing of the system turn (tool schemas), the user turn or the tool results is
|
||||
"""
|
||||
import argparse
|
||||
import importlib
|
||||
import re
|
||||
import json
|
||||
import os
|
||||
import random
|
||||
@@ -33,6 +34,55 @@ from transformers import AutoTokenizer # noqa: E402
|
||||
POOL = os.path.join(ROOT, "tasks_gen", "train")
|
||||
|
||||
|
||||
START_PREFIX = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
|
||||
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
|
||||
LIST_TOOLS = ("sap_inactive_objects", "sap_search_object", "sap_usage_references")
|
||||
READ_TOOLS = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
|
||||
"sap_syntax_check", "sap_atc_run")
|
||||
|
||||
|
||||
def foreign_name(name, own):
|
||||
n = (name or "").upper()
|
||||
m = START_PREFIX.match(n) or MID_PREFIX.match(n)
|
||||
return bool(m) and m.group(1).upper() != own.upper()
|
||||
|
||||
|
||||
def scrub_foreign(msgs, own_prefix):
|
||||
"""Remove other runs' leftover objects from list results (the proxy hides them since 2026-10-06; the older records still have them).
|
||||
Returns (messages, removed list entries, number of reads of a foreign object)."""
|
||||
tool_of = {}
|
||||
for m in msgs:
|
||||
if m["role"] == "assistant":
|
||||
for c in m.get("tool_calls") or []:
|
||||
a = c["function"].get("arguments") or "{}"
|
||||
try:
|
||||
a = json.loads(a) if isinstance(a, str) else a
|
||||
except ValueError:
|
||||
a = {}
|
||||
tool_of[c.get("id")] = (c["function"]["name"], a)
|
||||
removed = reads = 0
|
||||
out = []
|
||||
for m in msgs:
|
||||
if m["role"] == "tool":
|
||||
name, args = tool_of.get(m.get("tool_call_id"), ("", {}))
|
||||
if name in READ_TOOLS and foreign_name(args.get("objectName"), own_prefix):
|
||||
reads += 1
|
||||
if name in LIST_TOOLS:
|
||||
body = m["content"]
|
||||
prefix = "ERROR: " if body.startswith("ERROR: ") else ""
|
||||
try:
|
||||
data = json.loads(body[len(prefix):])
|
||||
except ValueError:
|
||||
data = None
|
||||
if isinstance(data, list):
|
||||
keep = [d for d in data if not (isinstance(d, dict) and foreign_name(d.get("name"), own_prefix))]
|
||||
if len(keep) != len(data):
|
||||
removed += len(data) - len(keep)
|
||||
m = dict(m, content=prefix + json.dumps(keep))
|
||||
out.append(m)
|
||||
return out, removed, reads
|
||||
|
||||
|
||||
def pct(v, p):
|
||||
v = sorted(v)
|
||||
return v[min(len(v) - 1, int(p * len(v)))] if v else None
|
||||
@@ -65,6 +115,8 @@ def main():
|
||||
ap.add_argument("--min-score", type=float, default=80)
|
||||
ap.add_argument("--valid-frac", type=float, default=0.10)
|
||||
ap.add_argument("--seed", type=int, default=20261006)
|
||||
ap.add_argument("--no-scrub", action="store_true", help="keep other runs' leftover objects in list results (default: removed)")
|
||||
ap.add_argument("--keep-foreign-reads", action="store_true", help="keep trajectories in which the model read another run's object (default: dropped)")
|
||||
ap.add_argument("--hook", help="module:function; function(row) -> None to drop, or a float weight (repeat factor), or a dict "
|
||||
"{'keep': bool, 'weight': float, 'extra': {...}} (for example the own-test mutation score)")
|
||||
a = ap.parse_args()
|
||||
@@ -90,7 +142,14 @@ def main():
|
||||
if not ok:
|
||||
continue
|
||||
msgs, removed = acc.trim_loops(rec["messages"])
|
||||
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs)})
|
||||
scrubbed = freads = 0
|
||||
if not a.no_scrub:
|
||||
msgs, scrubbed, freads = scrub_foreign(msgs, rec["prefix"])
|
||||
if freads and not a.keep_foreign_reads:
|
||||
kind0 = mix.kind_of_task_dir(r["task"])
|
||||
report["dropped"][kind0]["read_of_another_runs_object"] += 1
|
||||
continue
|
||||
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs), "scrubbed": scrubbed, "freads": freads})
|
||||
report["steps"]["accepted_trajectories"] = len(cand)
|
||||
|
||||
# 2 eval overlap at build time (the task against every eval task)
|
||||
@@ -127,7 +186,7 @@ def main():
|
||||
t = c["rec"]["task"]
|
||||
samples.append({"id": f"{c['row']['task']}_r{c['row']['run']}", "task": c["row"]["task"], "family": family_of(c["row"]["task"]),
|
||||
"kind": c["kind"], "category": t.get("category"), "object_type": t.get("object_type"),
|
||||
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"],
|
||||
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"], "foreign_entries_removed": c["scrubbed"],
|
||||
"score": c["rec"]["metadata"].get("score"), "teacher": c["rec"]["teacher"], "weight": 1.0,
|
||||
"n_tokens": n, "assistant_tokens": nl, "text": text, "assistant_spans": spans})
|
||||
report["steps"]["after_token_limit"] = len(samples)
|
||||
@@ -250,6 +309,7 @@ Samples over {r['settings']['max_tokens']} tokens are dropped, never cut. CLAS i
|
||||
|
||||
## Known limits
|
||||
- Small data (see the table); kinds marked short have fewer than the minimum samples.
|
||||
- Until 2026-10-06 the test system still held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes). Their names are removed from list results (`sap_inactive_objects`, searches) in these samples; trajectories in which the model read such an object are dropped.
|
||||
- The proxy added local abaplint messages (`syntaxCheck`) to some failed writes; the EPOD server does not do this itself.
|
||||
- The teacher sees only the tool results of its own runs; a run that repaired an error is kept with its error (60 to 65 % of the samples).
|
||||
- Validation split is by task family; do not mix valid samples into the training set.
|
||||
|
||||
Reference in New Issue
Block a user