stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,51 +1,57 @@
|
|||||||
# Stage 2 training set, first build (2026-10-06)
|
# Stage 2 training set, build of 2026-10-06 06:15:28
|
||||||
|
|
||||||
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`). Hook for item D: `train/hooks_example.py`.
|
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
||||||
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||||
|
|
||||||
## Result
|
## Result
|
||||||
90 accepted DeepSeek trajectories -> 90 after the eval overlap check (0 removed) -> 83 after the 48k limit
|
81 accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> 81 after the eval overlap check
|
||||||
(**7 dropped over 48000 tokens: 6 DDLS, 1 INTF**; none cut) -> CLAS cap 35 %: 18 CLAS kept, **31 CLAS in reserve** (`stage2_reserve.jsonl`)
|
-> 75 after the 48k limit (none cut) -> CLAS cap 35 %: 14 CLAS kept, **35 CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||||
-> **47 train + 5 valid samples** (1.28 M + 0.12 M tokens, 0.41 M loss tokens in train, p50 24771, p95 41829, max 47694).
|
-> **36 train + 4 valid samples** (0.94 M + 0.09 M tokens, 0.32 M loss tokens in train, p50 25079, p95 41724, max 45118, repair share 0.86).
|
||||||
Loss mask checked on all 47 train samples: no span contains a tool result, the system turn or a user turn; loss share 32 % of the tokens; no token straddles a span boundary.
|
Dropped:
|
||||||
|
- FUNC: {'read_of_another_runs_object': 3}
|
||||||
|
- DDLS: {'read_of_another_runs_object': 5, 'over_48k': 5}
|
||||||
|
- TABL: {'read_of_another_runs_object': 1}
|
||||||
|
- INTF: {'over_48k': 1}
|
||||||
|
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
||||||
|
|
||||||
## Not good enough yet (kinds under the minimum of 25)
|
## Not good enough yet (kinds under the minimum of 25)
|
||||||
| kind | samples | families | short |
|
| kind | samples | families | short |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| CLAS | 18 | 18 | 7 |
|
| CLAS | 14 | 14 | 11 |
|
||||||
| INTF | 2 | 2 | 23 |
|
| INTF | 2 | 2 | 23 |
|
||||||
| DDLS | 13 | 13 | 12 |
|
| DDLS | 9 | 9 | 16 |
|
||||||
| FUNC | 11 | 11 | 14 |
|
| FUNC | 8 | 8 | 17 |
|
||||||
| PROG | 5 | 5 | 20 |
|
| PROG | 5 | 5 | 20 |
|
||||||
| TABL | 3 | 3 | 22 |
|
| TABL | 2 | 2 | 23 |
|
||||||
| STRU | 0 | 0 | 25 |
|
| STRU | 0 | 0 | 25 |
|
||||||
| MSAG | 0 | 0 | 25 |
|
| MSAG | 0 | 0 | 25 |
|
||||||
| EXC | 0 | 0 | 25 |
|
| EXC | 0 | 0 | 25 |
|
||||||
|
|
||||||
- **STRU, MSAG, exception: 0 trajectories**; INTF 2, TABL 3, PROG 5. Only CLAS, DDLS and FUNC are near their share. Nothing is filled with copies. This is what the restart plan (new kinds first) has to fix.
|
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
|
||||||
- **DDLS loses 6 of 19 trajectories to the 48k limit.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit. Options: raise the limit to 64k (memory test decides), or generate CDS tasks with shorter runs. Not decided.
|
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
|
||||||
- Validation: 5 samples ({'CLAS': 2, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
|
||||||
- 2 trajectories of one task and K variants are on the same side (family key); there are no same-task pairs yet in practice.
|
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
|
||||||
|
- Validation: 4 samples ({'CLAS': 1, 'FUNC': 1, 'DDLS': 1, 'PROG': 1}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||||
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||||
|
|
||||||
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||||
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 47 samples, 0.41 M loss tokens per epoch. Loss tokens per option (the table uses today's data):
|
Stage 1 train: 374 documents, 2.53 M tokens (all tokens carry loss). Stage 2 train today: 36 samples, 0.32 M loss tokens per epoch. Loss tokens per option (today's data):
|
||||||
|
|
||||||
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||||
|---|---|---|---|---|---|
|
|---|---|---|---|---|---|
|
||||||
| 2 | 3 | 5.05 M | 1.23 M | 20 % | 6.3 M |
|
| 2 | 3 | 5.05 M | 0.95 M | 16 % | 6.0 M |
|
||||||
| 1 | 3 | 2.53 M | 1.23 M | 33 % | 3.8 M |
|
| 1 | 3 | 2.53 M | 0.95 M | 27 % | 3.5 M |
|
||||||
| 0.5 | 3 | 1.26 M | 1.23 M | 49 % | 2.5 M |
|
| 0.5 | 3 | 1.26 M | 0.95 M | 43 % | 2.2 M |
|
||||||
| 0.33 | 3 | 0.83 M | 1.23 M | 60 % | 2.1 M |
|
| 0.33 | 3 | 0.83 M | 0.95 M | 53 % | 1.8 M |
|
||||||
| 1 | 3 | 2.53 M | 3.70 M | 59 % | 6.2 M |
|
| 1 | 3 | 2.53 M | 2.85 M | 53 % | 5.4 M |
|
||||||
|
|
||||||
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: if stage 2 had three times today's data, as expected after the restart.)
|
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
|
||||||
|
|
||||||
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
|
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
|
||||||
Reasons: (1) stage 2 is the behavior we want (repair after the first error, write after a few reads, the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
|
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
|
||||||
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against 1.2 M of stage 2: the model would mostly learn documents again, the format of the tool calls would be a small part.
|
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||||
(3) With 47 samples, more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss (5 samples today) is too thin to catch it, so keep an eye on the train loss curve and use the checkpoints.
|
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||||
(4) The rule scales with the data: today it means about 0.33 epochs of stage 1 (about 125 documents), with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||||
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||||
|
|
||||||
## Also built
|
## Also built
|
||||||
|
|||||||
70
train/build_doc.py
Normal file
70
train/build_doc.py
Normal file
@@ -0,0 +1,70 @@
|
|||||||
|
"""Writes docs/stage2-build.md from runs/stage2_data/build_report.json (run after train/build_stage2.py)."""
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
rep = json.load(open(os.path.join(ROOT, "runs", "stage2_data", "build_report.json")))
|
||||||
|
tr, va, rs = rep["train"], rep["valid"], rep["reserve"]
|
||||||
|
s1 = rep["stage1_stats"]["train"]["tokens"]
|
||||||
|
s1d = rep["stage1_stats"]["train"]["docs"]
|
||||||
|
L2 = tr["loss_tokens"]
|
||||||
|
|
||||||
|
|
||||||
|
def row(e1, e2, L=L2):
|
||||||
|
a, b = s1 * e1, L * e2
|
||||||
|
return f"| {e1} | {e2} | {a / 1e6:.2f} M | {b / 1e6:.2f} M | {100 * b / (a + b):.0f} % | {(a + b) / 1e6:.1f} M |"
|
||||||
|
|
||||||
|
|
||||||
|
drop = rep["dropped"]
|
||||||
|
dropped_lines = "\n".join(f"- {k}: {v}" for k, v in drop.items())
|
||||||
|
kinds = "\n".join(f"| {k} | {v['samples']} | {v['families']} | {v['short']} |" for k, v in rep["kinds"].items())
|
||||||
|
doc = f"""# Stage 2 training set, build of {rep['built']}
|
||||||
|
|
||||||
|
Builder: `train/build_stage2.py` (output `runs/stage2_data/`, data card `README.md`, report `build_report.json`; this page: `train/build_doc.py`). Hook for item D: `train/hooks_example.py`.
|
||||||
|
HF dataset (private): `erhankeseli/abap-stage2-data`. The local Qwen trajectories (series A) are not read.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
{rep['steps']['accepted_trajectories']} accepted DeepSeek trajectories (after the scrub of other runs' leftover objects and the drop of trajectories that read such an object) -> {rep['steps']['after_eval_overlap']} after the eval overlap check
|
||||||
|
-> {rep['steps']['after_token_limit']} after the 48k limit (none cut) -> CLAS cap 35 %: {rep['steps']['clas_cap']['clas_kept']} CLAS kept, **{rep['steps']['clas_cap']['clas_reserve']} CLAS in reserve** (`stage2_reserve.jsonl`)
|
||||||
|
-> **{tr['samples']} train + {va['samples']} valid samples** ({tr['tokens'] / 1e6:.2f} M + {va['tokens'] / 1e6:.2f} M tokens, {tr['loss_tokens'] / 1e6:.2f} M loss tokens in train, p50 {tr['p50']}, p95 {tr['p95']}, max {tr['max']}, repair share {tr['repair_share']}).
|
||||||
|
Dropped:
|
||||||
|
{dropped_lines}
|
||||||
|
Loss mask checked on every train sample: no span contains a tool result, the system turn or a user turn; loss share about 32 % of the tokens; no token straddles a span boundary.
|
||||||
|
|
||||||
|
## Not good enough yet (kinds under the minimum of {rep['settings']['min_per_kind']})
|
||||||
|
| kind | samples | families | short |
|
||||||
|
|---|---|---|---|
|
||||||
|
{kinds}
|
||||||
|
|
||||||
|
- **STRU, MSAG, exception: 0 trajectories**; INTF, TABL, PROG only a few. The data is CLAS, DDLS and FUNC. Nothing is filled with copies; the restart plan (new kinds first) has to fix this.
|
||||||
|
- **DDLS loses trajectories to the 48k limit and to the drop of reads of other runs' objects.** The long CDS trajectories (own CDS test class, many reads) are exactly the ones over the limit.
|
||||||
|
Options: raise the limit to 64k (the memory test decides), or generate CDS tasks with shorter runs. Not decided.
|
||||||
|
- **Reads of another run's object:** the test system held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes); the proxy hides them since 2026-10-06. In the older trajectories their names are removed from list results (scrub) and trajectories in which the model read such an object are dropped (`--keep-foreign-reads` keeps them).
|
||||||
|
- Validation: {va['samples']} samples ({va['by_kind']}); the new kinds have no validation sample. After the restart the valid set must be rebuilt.
|
||||||
|
- Eval overlap: no accepted task overlaps an eval task (spec cosine 0.75, rules 0.60, names 0.60).
|
||||||
|
|
||||||
|
## Stage 1 : stage 2 ratio (proposal, Kral decides)
|
||||||
|
Stage 1 train: {s1d} documents, {s1 / 1e6:.2f} M tokens (all tokens carry loss). Stage 2 train today: {tr['samples']} samples, {L2 / 1e6:.2f} M loss tokens per epoch. Loss tokens per option (today's data):
|
||||||
|
|
||||||
|
| stage 1 epochs | stage 2 epochs | stage 1 loss tokens | stage 2 loss tokens | stage 2 share | total tokens seen* |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
{row(2, 3)}
|
||||||
|
{row(1, 3)}
|
||||||
|
{row(0.5, 3)}
|
||||||
|
{row(0.33, 3)}
|
||||||
|
{row(1, 3, 3 * L2)}
|
||||||
|
|
||||||
|
(*stage 1 tokens plus all stage 2 tokens per epoch; last row: stage 2 with three times today's data, as expected after the restart.)
|
||||||
|
|
||||||
|
**Proposal: stage 2 for 3 epochs always, stage 1 so that its loss tokens are about two thirds of the stage 2 loss tokens (stage 2 = 60 % of the loss).**
|
||||||
|
Reasons: (1) stage 2 is the behavior we want (repair after the first error: the teacher does it in 98 % of the cases, Qwen in 43 %; write after a few reads; the new kinds), and its samples are long and rare; stage 1 is domain knowledge in document form and acts as a
|
||||||
|
regularizer, so it should not dominate the gradient. (2) The earlier plan of 2 epochs of stage 1 (748 steps) would give 5.1 M stage 1 loss tokens against about 1 M of stage 2: the model would mostly learn documents again.
|
||||||
|
(3) With a few dozen samples more than 3 to 4 epochs of stage 2 risks memorizing them; the valid loss is too thin to catch it, so watch the train loss curve and use the checkpoints.
|
||||||
|
(4) The rule scales with the data: with today's data it means about 0.3 epochs of stage 1, with three times the stage 2 data about 1 epoch. `--s1-epochs` takes fractions.
|
||||||
|
Decision needed: this rule (60 % of the loss on stage 2), or a fixed 1 epoch of stage 1.
|
||||||
|
|
||||||
|
## Also built
|
||||||
|
`train/hf_train_bf16.py` (bf16, loss mask, mixing, memory test), `docs/bf16-memory.md` (memory table by GPU, estimates), `train/hooks_example.py` (hook for item D, weights).
|
||||||
|
"""
|
||||||
|
open(os.path.join(ROOT, "docs", "stage2-build.md"), "w").write(doc)
|
||||||
|
print("written", len(doc))
|
||||||
@@ -14,6 +14,7 @@ nothing of the system turn (tool schemas), the user turn or the tool results is
|
|||||||
"""
|
"""
|
||||||
import argparse
|
import argparse
|
||||||
import importlib
|
import importlib
|
||||||
|
import re
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
import random
|
import random
|
||||||
@@ -33,6 +34,55 @@ from transformers import AutoTokenizer # noqa: E402
|
|||||||
POOL = os.path.join(ROOT, "tasks_gen", "train")
|
POOL = os.path.join(ROOT, "tasks_gen", "train")
|
||||||
|
|
||||||
|
|
||||||
|
START_PREFIX = re.compile(r"^(Z\d[0-9A-Z]{6}_)", re.I)
|
||||||
|
MID_PREFIX = re.compile(r"^[A-Z]{1,5}_(Z\d[0-9A-Z]{6}_)", re.I)
|
||||||
|
LIST_TOOLS = ("sap_inactive_objects", "sap_search_object", "sap_usage_references")
|
||||||
|
READ_TOOLS = ("sap_pull_source", "sap_object_structure", "sap_object_members", "sap_element_info", "sap_run_unit_test", "sap_check_object",
|
||||||
|
"sap_syntax_check", "sap_atc_run")
|
||||||
|
|
||||||
|
|
||||||
|
def foreign_name(name, own):
|
||||||
|
n = (name or "").upper()
|
||||||
|
m = START_PREFIX.match(n) or MID_PREFIX.match(n)
|
||||||
|
return bool(m) and m.group(1).upper() != own.upper()
|
||||||
|
|
||||||
|
|
||||||
|
def scrub_foreign(msgs, own_prefix):
|
||||||
|
"""Remove other runs' leftover objects from list results (the proxy hides them since 2026-10-06; the older records still have them).
|
||||||
|
Returns (messages, removed list entries, number of reads of a foreign object)."""
|
||||||
|
tool_of = {}
|
||||||
|
for m in msgs:
|
||||||
|
if m["role"] == "assistant":
|
||||||
|
for c in m.get("tool_calls") or []:
|
||||||
|
a = c["function"].get("arguments") or "{}"
|
||||||
|
try:
|
||||||
|
a = json.loads(a) if isinstance(a, str) else a
|
||||||
|
except ValueError:
|
||||||
|
a = {}
|
||||||
|
tool_of[c.get("id")] = (c["function"]["name"], a)
|
||||||
|
removed = reads = 0
|
||||||
|
out = []
|
||||||
|
for m in msgs:
|
||||||
|
if m["role"] == "tool":
|
||||||
|
name, args = tool_of.get(m.get("tool_call_id"), ("", {}))
|
||||||
|
if name in READ_TOOLS and foreign_name(args.get("objectName"), own_prefix):
|
||||||
|
reads += 1
|
||||||
|
if name in LIST_TOOLS:
|
||||||
|
body = m["content"]
|
||||||
|
prefix = "ERROR: " if body.startswith("ERROR: ") else ""
|
||||||
|
try:
|
||||||
|
data = json.loads(body[len(prefix):])
|
||||||
|
except ValueError:
|
||||||
|
data = None
|
||||||
|
if isinstance(data, list):
|
||||||
|
keep = [d for d in data if not (isinstance(d, dict) and foreign_name(d.get("name"), own_prefix))]
|
||||||
|
if len(keep) != len(data):
|
||||||
|
removed += len(data) - len(keep)
|
||||||
|
m = dict(m, content=prefix + json.dumps(keep))
|
||||||
|
out.append(m)
|
||||||
|
return out, removed, reads
|
||||||
|
|
||||||
|
|
||||||
def pct(v, p):
|
def pct(v, p):
|
||||||
v = sorted(v)
|
v = sorted(v)
|
||||||
return v[min(len(v) - 1, int(p * len(v)))] if v else None
|
return v[min(len(v) - 1, int(p * len(v)))] if v else None
|
||||||
@@ -65,6 +115,8 @@ def main():
|
|||||||
ap.add_argument("--min-score", type=float, default=80)
|
ap.add_argument("--min-score", type=float, default=80)
|
||||||
ap.add_argument("--valid-frac", type=float, default=0.10)
|
ap.add_argument("--valid-frac", type=float, default=0.10)
|
||||||
ap.add_argument("--seed", type=int, default=20261006)
|
ap.add_argument("--seed", type=int, default=20261006)
|
||||||
|
ap.add_argument("--no-scrub", action="store_true", help="keep other runs' leftover objects in list results (default: removed)")
|
||||||
|
ap.add_argument("--keep-foreign-reads", action="store_true", help="keep trajectories in which the model read another run's object (default: dropped)")
|
||||||
ap.add_argument("--hook", help="module:function; function(row) -> None to drop, or a float weight (repeat factor), or a dict "
|
ap.add_argument("--hook", help="module:function; function(row) -> None to drop, or a float weight (repeat factor), or a dict "
|
||||||
"{'keep': bool, 'weight': float, 'extra': {...}} (for example the own-test mutation score)")
|
"{'keep': bool, 'weight': float, 'extra': {...}} (for example the own-test mutation score)")
|
||||||
a = ap.parse_args()
|
a = ap.parse_args()
|
||||||
@@ -90,7 +142,14 @@ def main():
|
|||||||
if not ok:
|
if not ok:
|
||||||
continue
|
continue
|
||||||
msgs, removed = acc.trim_loops(rec["messages"])
|
msgs, removed = acc.trim_loops(rec["messages"])
|
||||||
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs)})
|
scrubbed = freads = 0
|
||||||
|
if not a.no_scrub:
|
||||||
|
msgs, scrubbed, freads = scrub_foreign(msgs, rec["prefix"])
|
||||||
|
if freads and not a.keep_foreign_reads:
|
||||||
|
kind0 = mix.kind_of_task_dir(r["task"])
|
||||||
|
report["dropped"][kind0]["read_of_another_runs_object"] += 1
|
||||||
|
continue
|
||||||
|
cand.append({"rec": rec, "msgs": msgs, "removed": removed, "row": r, "repair": acc.has_repair(msgs), "scrubbed": scrubbed, "freads": freads})
|
||||||
report["steps"]["accepted_trajectories"] = len(cand)
|
report["steps"]["accepted_trajectories"] = len(cand)
|
||||||
|
|
||||||
# 2 eval overlap at build time (the task against every eval task)
|
# 2 eval overlap at build time (the task against every eval task)
|
||||||
@@ -127,7 +186,7 @@ def main():
|
|||||||
t = c["rec"]["task"]
|
t = c["rec"]["task"]
|
||||||
samples.append({"id": f"{c['row']['task']}_r{c['row']['run']}", "task": c["row"]["task"], "family": family_of(c["row"]["task"]),
|
samples.append({"id": f"{c['row']['task']}_r{c['row']['run']}", "task": c["row"]["task"], "family": family_of(c["row"]["task"]),
|
||||||
"kind": c["kind"], "category": t.get("category"), "object_type": t.get("object_type"),
|
"kind": c["kind"], "category": t.get("category"), "object_type": t.get("object_type"),
|
||||||
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"],
|
"attempt": c["row"]["attempt"], "repair": c["repair"], "loop_pairs_removed": c["removed"], "foreign_entries_removed": c["scrubbed"],
|
||||||
"score": c["rec"]["metadata"].get("score"), "teacher": c["rec"]["teacher"], "weight": 1.0,
|
"score": c["rec"]["metadata"].get("score"), "teacher": c["rec"]["teacher"], "weight": 1.0,
|
||||||
"n_tokens": n, "assistant_tokens": nl, "text": text, "assistant_spans": spans})
|
"n_tokens": n, "assistant_tokens": nl, "text": text, "assistant_spans": spans})
|
||||||
report["steps"]["after_token_limit"] = len(samples)
|
report["steps"]["after_token_limit"] = len(samples)
|
||||||
@@ -250,6 +309,7 @@ Samples over {r['settings']['max_tokens']} tokens are dropped, never cut. CLAS i
|
|||||||
|
|
||||||
## Known limits
|
## Known limits
|
||||||
- Small data (see the table); kinds marked short have fewer than the minimum samples.
|
- Small data (see the table); kinds marked short have fewer than the minimum samples.
|
||||||
|
- Until 2026-10-06 the test system still held leftover objects of earlier runs (the model's own `ZCL_<prefix>_...` classes). Their names are removed from list results (`sap_inactive_objects`, searches) in these samples; trajectories in which the model read such an object are dropped.
|
||||||
- The proxy added local abaplint messages (`syntaxCheck`) to some failed writes; the EPOD server does not do this itself.
|
- The proxy added local abaplint messages (`syntaxCheck`) to some failed writes; the EPOD server does not do this itself.
|
||||||
- The teacher sees only the tool results of its own runs; a run that repaired an error is kept with its error (60 to 65 % of the samples).
|
- The teacher sees only the tool results of its own runs; a run that repaired an error is kept with its error (60 to 65 % of the samples).
|
||||||
- Validation split is by task family; do not mix valid samples into the training set.
|
- Validation split is by task family; do not mix valid samples into the training set.
|
||||||
|
|||||||
Reference in New Issue
Block a user