Overlap limit for whole spec 0.75, slot logs; STATE and roadmap: stage 2 data progress

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 12:04:11 +02:00
parent 61b903f7b3
commit 35f0eb0e0c
4 changed files with 31 additions and 4 deletions

View File

@@ -120,3 +120,23 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- Tool names stay as in the EPOD ABAP MCP server (generic_v0 not used). The ADT "save failed" response cannot be fixed on the server side; the proxy adds a syntax check instead.
- HF cleanup done: test, overfit and full adapter repos deleted (step 200 and 400 checkpoints too). Kept: dataset `erhankeseli/abap-stage1-data`. Local adapter folders and `data_overfit` deleted.
- `BUDGET_LIMIT_USD` = 161 (ledger 66.97 + 35 usage x 2.7), until 12 October.
## Stage 2 data, Part B (2026-10-05)
- Teacher budget: `BUDGET_LIMIT_USD` 161 (ledger). Until 11 October up to 35 usage (= 94.5 ledger) without asking.
- B1 generator training mode: `harness/trainset.py`, pool `tasks_gen/train` (ids G1000+, run numbers 32000+), plan in
`tasks_gen/train/plan.json` (223 slots: category shares as the eval plan, 30 error tasks: named-type 10, reserved-word 10,
long-names 10), overlap check against all eval tasks and earlier training tasks (`harness/overlap.py`: spec cosine 0.75,
Goal + Business rules cosine 0.60, contract name Jaccard 0.60; calibrated on the eval pairs; K variants are the only
eval pairs above 0.75). A too close bundle goes back to the model as a repair message. K variants are not made for
training (they need the generic_v0 schema, which is not used). Started 2026-10-05 12:00 with 3 workers, target 200
accepted tasks, phase limit 27 ledger USD; logs `runs/gen_train/w*.log`, slot logs `tasks_gen/train/_logs/`.
- B2a proxy: a write that fails with only "An error occured during the save operation" gets `syntaxCheck` messages from local
abaplint parser (line + text). The EPOD syntax check cannot check the rejected source (tested 2026-10-05: it returned
0 errors for the stored stub), so abaplint is used. The raw result stays in `trajectory.jsonl` (`raw_result`).
- B2b record: `harness/record.py` writes `record.json` (task metadata, exact messages, tool schemas, raw tool results,
metadata: score, cost, end reason, ...) and `reasoning.json` (teacher reasoning, not training data) per run.
- B2c converter: `train/to_qwen.py` (Qwen 3.8 chat template, `enable_thinking=False`, Qwen XML tool call format; tokenizer
round-trip check of every tool call; assistant spans for loss masking). B2d filter: `train/accept.py`.
- Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.