stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-06 06:16:02 +02:00
parent c2c3803b1b
commit c5f1a36df0
3 changed files with 163 additions and 27 deletions

View File

@@ -1,51 +1,57 @@
# Stage 2 training set, first build (2026-10-06)
# Stage 2 training set, build of 2026-10-06 06:15:28
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`). Hook for item D: `train/hooks_example.py`.
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
## Result
90 accepted DeepSeek trajectories -> 90 after the eval overlap check (0 removed) -> 83 after the 48k limit
(**7 dropped over 48000 tokens: 6 DDLS, 1 INTF**; none cut) -> CLAS cap 35 %: 18 CLAS kept, **31 CLAS in reserve** (`stage2_reserve.jsonl`)
-> **47 train + 5 valid samples** (1.28 M + 0.12 M tokens, 0.41 M loss tokens in train, p50 24771, p95 41829, max 47694).
Loss mask checked on all 47 train samples: no span contains a tool result, the system turn or a user turn; loss share 32 % of the tokens; no token straddles a span boundary.
81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
Dropped:
- FUNC: {'read_of_another_runs_object': 3}
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
- TABL: {'read_of_another_runs_object': 1}
- INTF: {'over_48k': 1}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
## Not good enough yet (kinds under the minimum of 25)
| kind | samples | families | short |
|---|---|---|---|
| CLAS | 18 | 18 | 7 |
| CLAS | 14 | 14 | 11 |
| INTF | 2 | 2 | 23 |
| DDLS | 13 | 13 | 12 |
| FUNC | 11 | 11 | 14 |
| DDLS | 9 | 9 | 16 |
| FUNC | 8 | 8 | 17 |
| PROG | 5 | 5 | 20 |
| TABL | 3 | 3 | 22 |
| TABL | 2 | 2 | 23 |
| STRU | 0 | 0 | 25 |
| MSAG | 0 | 0 | 25 |
| EXC | 0 | 0 | 25 |
- **STRU, MSAG, exception: 0 trajectories**; INTF 2, TABL 3, PROG 5. Only CLAS, DDLS and FUNC are near their share. Nothing is filled with copies. This is what the restart plan (new kinds first) has to fix.
- **DDLS loses 6 of 19 trajectories to the 48k limit.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (memory test decides), or generate CDS tasks with shorter runs. Not decided.
- Validation: 5 samples ({'CLAS': 2, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- 2 trajectories of one task and K variants are on the same side (family key); there are no same-task pairs yet in practice.
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides)
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 47 samples, 0.41 M loss tokens per epoch. Loss tokens per option (the table uses today's data):
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
| 2 | 3 | 5.05 M | 1.23 M | 20 % | 6.3 M |
| 1 | 3 | 2.53 M | 1.23 M | 33 % | 3.8 M |
| 0.5 | 3 | 1.26 M | 1.23 M | 49 % | 2.5 M |
| 0.33 | 3 | 0.83 M | 1.23 M | 60 % | 2.1 M |
| 1 | 3 | 2.53 M | 3.70 M | 59 % | 6.2 M |
| 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M |
| 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M |
| 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M |
| 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M |
| 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M |
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: if stage 2 had three times today's data, as expected after the restart.)
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
Reasons: (1) stage 2 is the behavior we want (repair after the first error, write after a few reads, the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against 1.2 M of stage 2: the model would mostly learn documents again, the format of the tool calls would be a small part.
(3) With 47 samples, more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss (5 samples today) is too thin to catch it, so keep an eye on the train loss curve and use the checkpoints.
(4) The rule scales with the data: today it means about 0.33 epochs of stage 1 (about 125 documents), with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
## Also built

70
train/build_doc.py Normal file
View File

@@ -0,0 +1,70 @@
"""Writes docs/stage2-build.md from runs/stage2_data/build_report.json (run after train/build_stage2.py)."""
import json
import os
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
rep = json.load(open(os.path.join(ROOT, "runs", "stage2_data", "build_report.json")))
tr, va, rs = rep["train"], rep["valid"], rep["reserve"]
s1 = rep["stage1_stats"]["train"]["tokens"]
s1d = rep["stage1_stats"]["train"]["docs"]
L2 = tr["loss_tokens"]
def row(e1, e2, L=L2):
a, b = s1 * e1, L * e2
return f"| {e1} | {e2} | {a / 1e6:.2f} M | {b / 1e6:.2f} M | {100 * b / (a + b):.0f} % | {(a + b) / 1e6:.1f} M |"
drop = rep["dropped"]
dropped_lines = "\n".join(f"- {k}: {v}" for k, v in drop.items())
kinds = "\n".join(f"| {k} | {v['samples']} | {v['families']} | {v['short']} |" for k, v in rep["kinds"].items())
doc = f"""# Stage 2 training set, build of {rep['built']}
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
## Result
{rep['steps']['accepted_trajectories']} accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> {rep['steps']['after_eval_overlap']} after the eval overlap check
-> {rep['steps']['after_token_limit']} after the 48k limit (none cut) -> CLAS cap 35 %: {rep['steps']['clas_cap']['clas_kept']} CLAS kept, **{rep['steps']['clas_cap']['clas_reserve']} CLAS in reserve** (`stage2_reserve.jsonl`)
-> **{tr['samples']} train + {va['samples']} valid samples** ({tr['tokens'] / 1e6:.2f} M + {va['tokens'] / 1e6:.2f} M tokens, {tr['loss_tokens'] / 1e6:.2f} M loss tokens in train, p50 {tr['p50']}, p95 {tr['p95']}, max {tr['max']}, repair share {tr['repair_share']}).
Dropped:
{dropped_lines}
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
## Not good enough yet (kinds under the minimum of {rep['settings']['min_per_kind']})
| kind | samples | families | short |
|---|---|---|---|
{kinds}
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
## Stage 1 : stage 2 ratio (proposal, Kral decides)
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|---|---|---|---|---|---|
{row(2, 3)}
{row(1, 3)}
{row(0.5, 3)}
{row(0.33, 3)}
{row(1, 3, 3 * L2)}
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
## Also built
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
"""
open(os.path.join(ROOT, "docs", "stage2-build.md"), "w").write(doc)
print("written", len(doc))

View File

@@ -14,6 +14,7 @@ nothing of the system turn (tool schemas), the user turn or the tool results is
"""
import argparse
import importlib
import re
import json
import os
import random
@@ -33,6 +34,55 @@ from transformers import AutoTokenizer # noqa: E402
POOL = os.path.join(ROOT, "tasks_gen", "train")
START_PREFIX = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
LIST_TOOLS = ("sap_inactive_objects", "sap_search_object", "sap_usage_references")
READ_TOOLS = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
"sap_syntax_check", "sap_atc_run")
def foreign_name(name, own):
n = (name or "").upper()
m = START_PREFIX.match(n) or MID_PREFIX.match(n)
return bool(m) and m.group(1).upper() != own.upper()
def scrub_foreign(msgs, own_prefix):
"""Remove other runs' leftover objects from list results (the proxy hides them since 2026-10-06; the older records still have them).
Returns (messages, removed list entries, number of reads of a foreign object)."""
tool_of = {}
for m in msgs:
if m["role"] == "assistant":
for c in m.get("tool_calls") or []:
a = c["function"].get("arguments") or "{}"
try:
a = json.loads(a) if isinstance(a, str) else a
except ValueError:
a = {}
tool_of[c.get("id")] = (c["function"]["name"], a)
removed = reads = 0
out = []
for m in msgs:
if m["role"] == "tool":
name, args = tool_of.get(m.get("tool_call_id"), ("", {}))
if name in READ_TOOLS and foreign_name(args.get("objectName"), own_prefix):
reads += 1
if name in LIST_TOOLS:
body = m["content"]
prefix = "ERROR: " if body.startswith("ERROR: ") else ""
try:
data = json.loads(body[len(prefix):])
except ValueError:
data = None
if isinstance(data, list):
keep = [d for d in data if not (isinstance(d, dict) and foreign_name(d.get("name"), own_prefix))]
if len(keep) != len(data):
removed += len(data) - len(keep)
m = dict(m, content=prefix + json.dumps(keep))
out.append(m)
return out, removed, reads
def pct(v, p):
v = sorted(v)
return v[min(len(v) - 1, int(p * len(v)))] if v else None
@@ -65,6 +115,8 @@ def main():
ap.add_argument("--min-score", type=float, default=80)
ap.add_argument("--valid-frac", type=float, default=0.10)
ap.add_argument("--seed", type=int, default=20261006)
ap.add_argument("--no-scrub", action="store_true", help="keep other runs' leftover objects in list results (default: removed)")
ap.add_argument("--keep-foreign-reads", action="store_true", help="keep trajectories in which the model read another run's object (default: dropped)")
ap.add_argument("--hook", help="module:function; function(row) -> None to drop, or a float weight (repeat factor), or a dict "
"{'keep': bool, 'weight': float, 'extra': {...}} (for example the own-test mutation score)")
a = ap.parse_args()
@@ -90,7 +142,14 @@ def main():
if not ok:
continue
msgs, removed = acc.trim_loops(rec["messages"])
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs)})
scrubbed = freads = 0
if not a.no_scrub:
msgs, scrubbed, freads = scrub_foreign(msgs, rec["prefix"])
if freads and not a.keep_foreign_reads:
kind0 = mix.kind_of_task_dir(r["task"])
report["dropped"][kind0]["read_of_another_runs_object"] += 1
continue
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs), "scrubbed": scrubbed, "freads": freads})
report["steps"]["accepted_trajectories"] = len(cand)
# 2 eval overlap at build time (the task against every eval task)
@@ -127,7 +186,7 @@ def main():
t = c["rec"]["task"]
samples.append({"id": f"{c['row']['task']}_r{c['row']['run']}", "task": c["row"]["task"], "family": family_of(c["row"]["task"]),
"kind": c["kind"], "category": t.get("category"), "object_type": t.get("object_type"),
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"],
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"], "foreign_entries_removed": c["scrubbed"],
"score": c["rec"]["metadata"].get("score"), "teacher": c["rec"]["teacher"], "weight": 1.0,
"n_tokens": n, "assistant_tokens": nl, "text": text, "assistant_spans": spans})
report["steps"]["after_token_limit"] = len(samples)
@@ -250,6 +309,7 @@ Samples over {r['settings']['max_tokens']} tokens are dropped, never cut. CLAS i
## Known limits
- Small data (see the table); kinds marked short have fewer than the minimum samples.
- Until 2026-10-06 the test system still held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes). Their names are removed from list results (`sap_inactive_objects`, searches) in these samples; trajectories in which the model read such an object are dropped.
- The proxy added local abaplint messages (`syntaxCheck`) to some failed writes; the EPOD server does not do this itself.
- The teacher sees only the tool results of its own runs; a run that repaired an error is kept with its error (60 to 65 % of the samples).
- Validation split is by task family; do not mix valid samples into the training set.