- generator: budget floor (2x oracle activations, 3x calls), static checks for CDS $parameters and UNION annotation - harness/mutation.py: deterministic mutants of the reference; hidden tests must fail - G0022 revalidated with harness fixes: oracle 100, null 0 - results in docs/faz1-tasarim.md 11f Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
7.5 KiB
7.5 KiB
ABAP LLM harness — instructions for Claude Code
0. Working with Kral
- Talk in Turkish. Short answers. Remove words and sentences that add no value.
- English text (specs, prompts, docs for the model): ASD-STE100 Simplified Technical English.
- Ask before an action that costs much (cloud budget) or that changes the A4H system outside harness runs.
- For a series of updates to one document, deliver the result directly as a Markdown file; do not ask each time.
- Record each decision with its date in
docs/yol-haritasi.mdordocs/faz1-tasarim.md. - Commit the current changes before a new generation or test batch.
- Kral has no time for a full review. Eval review = automatic gates + empirical filter (2–3 models) + Claude review with a checklist; Kral reviews only flagged tasks + one per category (~10 tasks).
1. Project
- Goal: train an open-weight ABAP model (Apache 2.0). Base model: Apache 2.0 or MIT only.
- Role of the model: technical ABAP consultant. It writes ABAP, knows Clean ABAP and what to use how. No SAP module knowledge; the functional side (spec or the /sapplan planner) gives it.
- Contract of the model: "ABAP task as text + generic ABAP MCP interface". The plan format (plan-agent.md) gives the best result, but it is not mandatory.
- Training data: never Claude output. Eval tasks never go into training data.
- Input: the plan format gives the best result but is not mandatory. The old playbook format is not used.
- Generic ABAP MCP interface (core: search_object, read_source, object_structure, where_used, create_object, write_source, activate, syntax_check, run_unit_tests, run_atc). Small tool name/schema variation in training.
- Category K: incomplete or free-text input + other tool schema.
- Object type mix: CLAS/INTF 35 %, FUNC 15 %, PROG 10 %, DDIC 10 %, CDS 25 %, MSAG + exception 5 %.
- Teacher and task author: DeepSeek V4.1 Flash (Ollama cloud). Budget calendar in
docs/yol-haritasi.md. - Read first:
docs/yol-haritasi.md(roadmap, current position) anddocs/faz1-tasarim.md(design). When a decision or a result changes, update these files.
2. Environment (this Mac mini)
- A4H: Docker container
a4h, HTTPlocalhost:50000, client 001, SAP_BASIS 816 SP01. - MCP server (EPOD) runs inside ADT (Eclipse) at
127.0.0.1:3000. Token:.env(MCP_TOKEN). System nameA4H, mode write. - Ollama
127.0.0.1:11434:qwen3.8-27b-32k(local, num_ctx 32k, ~18 GB). Slow: 20–40 min per task.deepseek-v4.1-flash:cloud(Ollama cloud, MIT): teacher and task author. ~0.05–0.08 USD per run.
- Python 3.9 (system), Node; abaplint in
node_modules.
3. Rules
- Local model: one request at a time. No time limit (tool-call budget limits a run).
- Do not run local and cloud model runs at the same time: Ollama can queue them together.
- A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between
sessions; a parallel call gets "[LOCK] Concurrent call detected".
mcp_client.pyretries this. Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs. - Budget:
runs/ledger.jsonl(list prices, upper bound)..env:BUDGET_LIMIT_USD,BUDGET_CYCLE_START. Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work. - Objects: package
$TMPonly. PrefixZ+ run (4 chars base36) + task (3 chars base36) +_(harness/task.py). Teardown after each run with the ADT deletion API (adt_client.py, credentials in.env). Delete only objects with a run prefix. Clean up probe objects. - The model never sees delete/teardown. The proxy (
proxy.py) has a tool whitelist and hides other runs' objects. - To wait for a background job, use its PID:
while kill -0 <PID> 2>/dev/null; do sleep 60; done. Do NOT usepgrep -f <pattern>in the loop: it also finds the loop's own command line, so the loop never ends. The process name of a module run isPython -m harness.<module>(notpython3);pgrep -f "harness.pilot"without a loop gives the PID. - Start long jobs with
nohup ... &from a separate line aftercd(acd ... && nohup ... &chain runs in a subshell; the commands after it do not see the new directory).
4. Layout
harness/mcp_client.pyMCP client (retry on 404 and on LOCK).harness/adt_client.pyADT deletion API.harness/proxy.pytool whitelist, budget, prefix filter, trajectory log.harness/agents.pyoracle, null, llm (OpenAI-compatible; retries; ledger).harness/runner.pysetup → agent → gates G1–G6 → hidden tests → own tests → ATC → abaplint → score → teardown.harness/generator.pytask generator (cloud model writes a bundle; validation oracle = 100, null = 0; max 3 repairs).harness/pilot.pypilot list (20 tasks G0002–G0021).harness/ledger.pycost ledger and budget guard.harness/cli.pyrun, rescore, teardown, teardown-all, cleanup-list.tasks/hand-written tasks T01 (CLAS), T13 (FUNC), T14 (PROG + ALV), T15 (CDS). Oracle 100, null 0 for all.tasks_gen/eval/,tasks_gen/train/generated tasks (separate pools).runs/run results (git ignores it).runs/_archive_v1old runs.
5. Commands
python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N
python3 -m harness.cli rescore runs/<run_dir>
python3 -m harness.cli cleanup-list <PREFIX> runs/<dir> && python3 -m harness.cli teardown runs/<dir>
python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000
python3 -m harness.pilot 2 # pilot list; `python3 -m harness.pilot 22 hard 1400` hard batch
python3 -m harness.mutation G0002 --pool eval --run-base 5000 [--keep]
python3 -c "from harness.ledger import spent; print(spent())"
6. Current state (2026-10-02)
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support; task generator; ledger and budget guard.
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
- Pilot (G0002–G0021): 16/20 accepted, 0.16 USD per accepted task (
docs/faz1-tasarim.md11e). G0020 and G0013 fixed in review and revalidated → 18/20. - Hard batch (G0022–G0026, 2026-10-03): 5/5 accepted (G0022 after harness fixes), 0.097 USD per accepted
task, first-attempt 1/5 (
docs/faz1-tasarim.md11f). - Generator: static checks, local abaplint parser check, dependency order of seed/reference, budget floor, max_tokens 80k. Runner: G2 for CDS with parameters, G6 ignores unknown standard superclasses.
- Step D:
harness/mutation.py(deterministic mutants of the reference, hidden tests must fail).
7. Next steps
- Step D: run the mutation check on all accepted tasks; integrate it into the generator.
- Use killed mutants as faulty references for own-test scoring (component (a)).
- Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a checklist, Kral spot-checks ~10 flagged tasks.
- Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
- Step 1.3b: generic ABAP MCP interface (spec
abap-mcp-arayuz.md, later). The proxy is the first adapter.
8. Open items for Kral (EPOD server)
- When a write fails, run a syntax check and return its messages. Now the server says only "save failed"
(for example for
TYPE c LENGTH nin a method signature). The model cannot see the cause and cannot repair; this hurts eval and training. Workaround in the generator: local abaplint parser check. - Concurrency: queue calls per RFC connection, or use a connection pool.
- BDEF creation (needed for RAP tasks).