# ABAP LLM harness — instructions for Claude Code ## 0. Working with Kral - Talk in Turkish. Short answers. Remove words and sentences that add no value. - English text (specs, prompts, docs for the model): ASD-STE100 Simplified Technical English. - Ask before an action that costs much (cloud budget) or that changes the A4H system outside harness runs. - For a series of updates to one document, deliver the result directly as a Markdown file; do not ask each time. - Record each decision with its date in `docs/yol-haritasi.md` or `docs/faz1-tasarim.md`. - Commit the current changes before a new generation or test batch. - Kral has no time for a full review. Eval review = automatic gates + empirical filter (2–3 models) + Claude review with a checklist; Kral reviews only flagged tasks + one per category (~10 tasks). ## 1. Project - Goal: train an open-weight ABAP model (Apache 2.0). Base model: Apache 2.0 or MIT only. - Role of the model: technical ABAP consultant. It writes ABAP, knows Clean ABAP and what to use how. No SAP module knowledge; the functional side (spec or the /sapplan planner) gives it. - Contract of the model: "ABAP task as text + generic ABAP MCP interface". The plan format (plan-agent.md) gives the best result, but it is not mandatory. - Training data: never Claude output. Eval tasks never go into training data. - Input: the plan format gives the best result but is not mandatory. The old playbook format is not used. - Generic ABAP MCP interface (core: search_object, read_source, object_structure, where_used, create_object, write_source, activate, syntax_check, run_unit_tests, run_atc). Small tool name/schema variation in training. - Category K: incomplete or free-text input + other tool schema. - Object type mix: CLAS/INTF 35 %, FUNC 15 %, PROG 10 %, DDIC 10 %, CDS 25 %, MSAG + exception 5 %. - Teacher and task author: DeepSeek V4.1 Flash (Ollama cloud). Budget calendar in `docs/yol-haritasi.md`. - Read first: `docs/yol-haritasi.md` (roadmap, current position) and `docs/faz1-tasarim.md` (design). When a decision or a result changes, update these files. ## 2. Environment (this Mac mini) - A4H: Docker container `a4h`, HTTP `localhost:50000`, client 001, SAP_BASIS 816 SP01. - MCP server (EPOD) runs inside ADT (Eclipse) at `127.0.0.1:3000`. Token: `.env` (`MCP_TOKEN`). System name `A4H`, mode write. - Ollama `127.0.0.1:11434`: - `qwen3.8-27b-32k` (local, num_ctx 32k, ~18 GB). Slow: 20–40 min per task. - `deepseek-v4.1-flash:cloud` (Ollama cloud, MIT): teacher and task author. ~0.05–0.08 USD per run. - Python 3.9 (system), Node; abaplint in `node_modules`. ## 3. Rules - Local model: one request at a time. No time limit (tool-call budget limits a run). - Do not run local and cloud model runs at the same time: Ollama can queue them together. - A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this. Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs. - Budget: `runs/ledger.jsonl` (list prices, upper bound; real cost ≈ ledger / 2.7, measured 2026-10-03 20:15: ledger 66.97 vs Ollama monthly usage 24.63 USD (earlier ratio 1.47 was wrong)). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`. Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work. - Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_` (`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`). Delete only objects with a run prefix. Clean up probe objects. - The model never sees delete/teardown. The proxy (`proxy.py`) has a tool whitelist and hides other runs' objects. - To wait for a background job, use its PID: `while kill -0 2>/dev/null; do sleep 60; done`. Do NOT use `pgrep -f ` in the loop: it also finds the loop's own command line, so the loop never ends. The process name of a module run is `Python -m harness.` (not `python3`); `pgrep -f "harness.pilot"` without a loop gives the PID. - Start long jobs with `nohup ... &` from a separate line after `cd` (a `cd ... && nohup ... &` chain runs in a subshell; the commands after it do not see the new directory). ## 4. Layout - `harness/mcp_client.py` MCP client (retry on 404 and on LOCK). - `harness/adt_client.py` ADT deletion API. - `harness/proxy.py` tool whitelist, budget, prefix filter, trajectory log. - `harness/agents.py` oracle, null, llm (OpenAI-compatible; retries; ledger). - `harness/runner.py` setup → agent → gates G1–G6 → hidden tests → own tests → ATC → abaplint → score → teardown. - `harness/generator.py` task generator (cloud model writes a bundle; validation oracle = 100, null = 0; max 3 repairs). - `harness/pilot.py` pilot list (20 tasks G0002–G0021). - `harness/ledger.py` cost ledger and budget guard. - `harness/cli.py` run, rescore, teardown, teardown-all, cleanup-list. - `tasks/` hand-written tasks T01 (CLAS), T13 (FUNC), T14 (PROG + ALV), T15 (CDS). Oracle 100, null 0 for all. - `tasks_gen/eval/`, `tasks_gen/train/` generated tasks (separate pools). - `runs/` run results (git ignores it). `runs/_archive_v1` old runs. ## 5. Commands ``` python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N python3 -m harness.cli rescore runs/ python3 -m harness.cli cleanup-list runs/ && python3 -m harness.cli teardown runs/ python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000 python3 -m harness.pilot 2 # pilot list; `python3 -m harness.pilot 22 hard 1400` hard batch python3 -m harness.mutation G0002 --pool eval --run-base 5000 [--keep] python3 -c "from harness.ledger import spent; print(spent())" ``` ## 6. Current state (2026-10-02) - Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support; task generator; ledger and budget guard. - Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100. - Pilot (G0002–G0021): 16/20 accepted, 0.16 USD per accepted task (`docs/faz1-tasarim.md` 11e). G0020 and G0013 fixed in review and revalidated → 18/20. - Hard batch (G0022–G0026, 2026-10-03): 5/5 accepted (G0022 after harness fixes), 0.097 USD per accepted task, first-attempt 1/5 (`docs/faz1-tasarim.md` 11f). - Generator: static checks, local abaplint parser check, dependency order of seed/reference, budget floor, max_tokens 80k. Runner: G2 for CDS with parameters, G6 ignores unknown standard superclasses. - Step D done: `harness/mutation.py`, part of generator acceptance. All 26 tasks pass (`docs/faz1-tasarim.md` 11g). Killed mutants are kept in `/faulty/`. - Step 5 code: H stop scoring (`judge.py`), K variants (`make_k_variant`), proxy tool schema `generic_v0` (draft). Not yet run on SAP. - Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started. ## 6a. Base model decision (2026-10-04, revised) - **Base model: Qwen 3.8 27B** (`Qwen/Qwen3.8-27B`, Apache 2.0; weights `mlx-community/Qwen3.8-27B-4bit`, local `~/models/Qwen3.8-27B-4bit`). Kral decision, same day, after the Devstral baseline. - Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training must fix, and he does not like Devstral. Devstral is not used further. - Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0. `runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`. - A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference. - Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl). ## 7. Next steps 1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`). 2. Own-test scoring part (a): run the model's own tests against `faulty/` mutants. 3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a checklist, Kral spot-checks ~10 flagged tasks. 4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema). 5. Step 1.3b: generic ABAP MCP interface (spec `abap-mcp-arayuz.md`, later). The proxy is the first adapter. ## 8. Open items for Kral (EPOD server) - When a write fails, run a syntax check and return its messages. Now the server says only "save failed" (for example for `TYPE c LENGTH n` in a method signature). The model cannot see the cause and cannot repair; this hurts eval and training. Workaround in the generator: local abaplint parser check. - Unit test timeout: an endless loop in tested code blocks the one RFC connection for ~8 min (mutation run 5201, G0002). Request: stop a unit test run after N seconds (DURATION SHORT = 60 s). - Concurrency: queue calls per RFC connection, or use a connection pool. - BDEF creation (needed for RAP tasks).