Files
abap-llm/CLAUDE.md

9.4 KiB
Raw Blame History

ABAP LLM harness — instructions for Claude Code

0. Working with Kral

  • Talk in Turkish. Short answers. Remove words and sentences that add no value.
  • English text (specs, prompts, docs for the model): ASD-STE100 Simplified Technical English.
  • Ask before an action that costs much (cloud budget) or that changes the A4H system outside harness runs.
  • For a series of updates to one document, deliver the result directly as a Markdown file; do not ask each time.
  • Record each decision with its date in docs/yol-haritasi.md or docs/faz1-tasarim.md.
  • Commit the current changes before a new generation or test batch.
  • Kral has no time for a full review. Eval review = automatic gates + empirical filter (2–3 models) + Claude review with a checklist; Kral reviews only flagged tasks + one per category (~10 tasks).

1. Project

  • Goal: train an open-weight ABAP model (Apache 2.0). Base model: Apache 2.0 or MIT only.
  • Role of the model: technical ABAP consultant. It writes ABAP, knows Clean ABAP and what to use how. No SAP module knowledge; the functional side (spec or the /sapplan planner) gives it.
  • Contract of the model: "ABAP task as text + generic ABAP MCP interface". The plan format (plan-agent.md) gives the best result, but it is not mandatory.
  • Training data: never Claude output. Eval tasks never go into training data.
  • Input: the plan format gives the best result but is not mandatory. The old playbook format is not used.
  • Generic ABAP MCP interface (core: search_object, read_source, object_structure, where_used, create_object, write_source, activate, syntax_check, run_unit_tests, run_atc). Small tool name/schema variation in training.
  • Category K: incomplete or free-text input + other tool schema.
  • Object type mix: CLAS/INTF 35 %, FUNC 15 %, PROG 10 %, DDIC 10 %, CDS 25 %, MSAG + exception 5 %.
  • Teacher and task author: DeepSeek V4.1 Flash (Ollama cloud). Budget calendar in docs/yol-haritasi.md.
  • Read first: docs/yol-haritasi.md (roadmap, current position) and docs/faz1-tasarim.md (design). When a decision or a result changes, update these files.

2. Environment (this Mac mini)

  • A4H: Docker container a4h, HTTP localhost:50000, client 001, SAP_BASIS 816 SP01.
  • MCP server (EPOD) runs inside ADT (Eclipse) at 127.0.0.1:3000. Token: .env (MCP_TOKEN). System name A4H, mode write.
  • Ollama 127.0.0.1:11434:
    • qwen3.8-27b-32k (local, num_ctx 32k, ~18 GB). Slow: 20–40 min per task.
    • deepseek-v4.1-flash:cloud (Ollama cloud, MIT): teacher and task author. ~0.05–0.08 USD per run.
  • Python 3.9 (system), Node; abaplint in node_modules.

3. Rules

  • Local model: one request at a time. No time limit (tool-call budget limits a run).
  • Do not run local and cloud model runs at the same time: Ollama can queue them together.
  • A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between sessions; a parallel call gets "[LOCK] Concurrent call detected". mcp_client.py retries this. Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs.
  • Budget: runs/ledger.jsonl (list prices, upper bound; real cost ≈ ledger / 2.7, measured 2026-10-03 20:15: ledger 66.97 vs Ollama monthly usage 24.63 USD (earlier ratio 1.47 was wrong)). .env: BUDGET_LIMIT_USD, BUDGET_CYCLE_START. Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work.
  • Objects: package $TMP only. Prefix Z + run (4 chars base36) + task (3 chars base36) + _ (harness/task.py). Teardown after each run with the ADT deletion API (adt_client.py, credentials in .env). Delete only objects with a run prefix. Clean up probe objects.
  • The model never sees delete/teardown. The proxy (proxy.py) has a tool whitelist and hides other runs' objects.
  • To wait for a background job, use its PID: while kill -0 <PID> 2>/dev/null; do sleep 60; done. Do NOT use pgrep -f <pattern> in the loop: it also finds the loop's own command line, so the loop never ends. The process name of a module run is Python -m harness.<module> (not python3); pgrep -f "harness.pilot" without a loop gives the PID.
  • Start long jobs with nohup ... & from a separate line after cd (a cd ... && nohup ... & chain runs in a subshell; the commands after it do not see the new directory).

4. Layout

  • harness/mcp_client.py MCP client (retry on 404 and on LOCK).
  • harness/adt_client.py ADT deletion API.
  • harness/proxy.py tool whitelist, budget, prefix filter, trajectory log.
  • harness/agents.py oracle, null, llm (OpenAI-compatible; retries; ledger).
  • harness/runner.py setup → agent → gates G1–G6 → hidden tests → own tests → ATC → abaplint → score → teardown.
  • harness/generator.py task generator (cloud model writes a bundle; validation oracle = 100, null = 0; max 3 repairs).
  • harness/pilot.py pilot list (20 tasks G0002–G0021).
  • harness/ledger.py cost ledger and budget guard.
  • harness/cli.py run, rescore, teardown, teardown-all, cleanup-list.
  • tasks/ hand-written tasks T01 (CLAS), T13 (FUNC), T14 (PROG + ALV), T15 (CDS). Oracle 100, null 0 for all.
  • tasks_gen/eval/, tasks_gen/train/ generated tasks (separate pools).
  • runs/ run results (git ignores it). runs/_archive_v1 old runs.

5. Commands

python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N
python3 -m harness.cli rescore runs/<run_dir>
python3 -m harness.cli cleanup-list <PREFIX> runs/<dir> && python3 -m harness.cli teardown runs/<dir>
python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000
python3 -m harness.pilot 2            # pilot list; `python3 -m harness.pilot 22 hard 1400` hard batch
python3 -m harness.mutation G0002 --pool eval --run-base 5000 [--keep]
python3 -c "from harness.ledger import spent; print(spent())"

6. Current state (2026-10-02)

  • Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support; task generator; ledger and budget guard.
  • Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
  • Pilot (G0002–G0021): 16/20 accepted, 0.16 USD per accepted task (docs/faz1-tasarim.md 11e). G0020 and G0013 fixed in review and revalidated → 18/20.
  • Hard batch (G0022–G0026, 2026-10-03): 5/5 accepted (G0022 after harness fixes), 0.097 USD per accepted task, first-attempt 1/5 (docs/faz1-tasarim.md 11f).
  • Generator: static checks, local abaplint parser check, dependency order of seed/reference, budget floor, max_tokens 80k. Runner: G2 for CDS with parameters, G6 ignores unknown standard superclasses.
  • Step D done: harness/mutation.py, part of generator acceptance. All 26 tasks pass (docs/faz1-tasarim.md 11g). Killed mutants are kept in <task>/faulty/.
  • Step 5 code: H stop scoring (judge.py), K variants (make_k_variant), proxy tool schema generic_v0 (draft). Not yet run on SAP.
  • Step F: harness/evalset.py (88 slots G0100–G0187). Not started.

6a. Base model decision (2026-10-04, revised)

  • Base model: Qwen 3.8 27B (Qwen/Qwen3.8-27B, Apache 2.0; weights mlx-community/Qwen3.8-27B-4bit, local ~/models/Qwen3.8-27B-4bit). Kral decision, same day, after the Devstral baseline.
  • Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training must fix, and he does not like Devstral. Devstral is not used further.
  • Devstral Small 2 tested and dropped (same settings, 11 tasks): mean 6.8 vs Qwen 15.8; 1/11 vs 3/11 tasks above 0; end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: runs/archive/devstral/. The Devstral weights were deleted. No further base model tests.
  • Official Qwen baseline (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook, run base 22000): runs/stage1/baseline.json (= baseline_qwen.json). Comparison: docs/stage1-baseline.md. Older partial Qwen runs (Mac mini, aborted or thinking on): runs/archive/qwen38/, not comparable.
  • Training tool is not chosen. Mac: mlx_lm.lora (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).

7. Next steps

  1. Test one H slot (G0168) and one K slot (G0178); then run step F (python3 -m harness.evalset run).
  2. Own-test scoring part (a): run the model's own tests against faulty/ mutants.
  3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a checklist, Kral spot-checks ~10 flagged tasks.
  4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
  5. Step 1.3b: generic ABAP MCP interface (spec abap-mcp-arayuz.md, later). The proxy is the first adapter.

8. Open items for Kral (EPOD server)

  • When a write fails, run a syntax check and return its messages. Now the server says only "save failed" (for example for TYPE c LENGTH n in a method signature). The model cannot see the cause and cannot repair; this hurts eval and training. Workaround in the generator: local abaplint parser check.
  • Unit test timeout: an endless loop in tested code blocks the one RFC connection for ~8 min (mutation run 5201, G0002). Request: stop a unit test run after N seconds (DURATION SHORT = 60 s).
  • Concurrency: queue calls per RFC connection, or use a connection pool.
  • BDEF creation (needed for RAP tasks).