Files
abap-llm/CLAUDE.md

138 lines
9.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ABAP LLM harness — instructions for Claude Code
## 0. Working with Kral
- Talk in Turkish. Short answers. Remove words and sentences that add no value.
- English text (specs, prompts, docs for the model): ASD-STE100 Simplified Technical English.
- Ask before an action that costs much (cloud budget) or that changes the A4H system outside harness runs.
- For a series of updates to one document, deliver the result directly as a Markdown file; do not ask each time.
- Record each decision with its date in `docs/yol-haritasi.md` or `docs/faz1-tasarim.md`.
- Commit the current changes before a new generation or test batch.
- Kral has no time for a full review. Eval review = automatic gates + empirical filter (2–3 models) +
Claude review with a checklist; Kral reviews only flagged tasks + one per category (~10 tasks).
## 1. Project
- Goal: train an open-weight ABAP model (Apache 2.0). Base model: Apache 2.0 or MIT only.
- Role of the model: technical ABAP consultant. It writes ABAP, knows Clean ABAP and what to use how.
No SAP module knowledge; the functional side (spec or the /sapplan planner) gives it.
- Contract of the model: "ABAP task as text + generic ABAP MCP interface".
The plan format (plan-agent.md) gives the best result, but it is not mandatory.
- Training data: never Claude output. Eval tasks never go into training data.
- Input: the plan format gives the best result but is not mandatory. The old playbook format is not used.
- Generic ABAP MCP interface (core: search_object, read_source, object_structure, where_used, create_object,
write_source, activate, syntax_check, run_unit_tests, run_atc). Small tool name/schema variation in training.
- Category K: incomplete or free-text input + other tool schema.
- Object type mix: CLAS/INTF 35 %, FUNC 15 %, PROG 10 %, DDIC 10 %, CDS 25 %, MSAG + exception 5 %.
- Teacher and task author: DeepSeek V4.1 Flash (Ollama cloud). Budget calendar in `docs/yol-haritasi.md`.
- Read first: `docs/yol-haritasi.md` (roadmap, current position) and `docs/faz1-tasarim.md` (design).
When a decision or a result changes, update these files.
## 2. Environment (this Mac mini)
- A4H: Docker container `a4h`, HTTP `localhost:50000`, client 001, SAP_BASIS 816 SP01.
- MCP server (EPOD) runs inside ADT (Eclipse) at `127.0.0.1:3000`. Token: `.env` (`MCP_TOKEN`).
System name `A4H`, mode write.
- Ollama `127.0.0.1:11434`:
- `qwen3.8-27b-32k` (local, num_ctx 32k, ~18 GB). Slow: 20–40 min per task.
- `deepseek-v4.1-flash:cloud` (Ollama cloud, MIT): teacher and task author. ~0.05–0.08 USD per run.
- Python 3.9 (system), Node; abaplint in `node_modules`.
## 3. Rules
- Local model: one request at a time. No time limit (tool-call budget limits a run).
- Do not run local and cloud model runs at the same time: Ollama can queue them together.
- A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between
sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this.
Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs.
- Budget: `runs/ledger.jsonl` (list prices, upper bound; real cost ≈ ledger / 2.7, measured 2026-10-03 20:15: ledger 66.97 vs Ollama monthly usage 24.63 USD (earlier ratio 1.47 was wrong)). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`.
Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work.
- Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_`
(`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`).
Delete only objects with a run prefix. Clean up probe objects.
- The model never sees delete/teardown. The proxy (`proxy.py`) has a tool whitelist and hides other runs' objects.
- To wait for a background job, use its PID: `while kill -0 <PID> 2>/dev/null; do sleep 60; done`.
Do NOT use `pgrep -f <pattern>` in the loop: it also finds the loop's own command line, so the loop never ends.
The process name of a module run is `Python -m harness.<module>` (not `python3`); `pgrep -f "harness.pilot"`
without a loop gives the PID.
- Start long jobs with `nohup ... &` from a separate line after `cd` (a `cd ... && nohup ... &` chain runs
in a subshell; the commands after it do not see the new directory).
## 4. Layout
- `harness/mcp_client.py` MCP client (retry on 404 and on LOCK).
- `harness/adt_client.py` ADT deletion API.
- `harness/proxy.py` tool whitelist, budget, prefix filter, trajectory log.
- `harness/agents.py` oracle, null, llm (OpenAI-compatible; retries; ledger).
- `harness/runner.py` setup → agent → gates G1–G6 → hidden tests → own tests → ATC → abaplint → score → teardown.
- `harness/generator.py` task generator (cloud model writes a bundle; validation oracle = 100, null = 0; max 3 repairs).
- `harness/pilot.py` pilot list (20 tasks G0002–G0021).
- `harness/ledger.py` cost ledger and budget guard.
- `harness/cli.py` run, rescore, teardown, teardown-all, cleanup-list.
- `tasks/` hand-written tasks T01 (CLAS), T13 (FUNC), T14 (PROG + ALV), T15 (CDS). Oracle 100, null 0 for all.
- `tasks_gen/eval/`, `tasks_gen/train/` generated tasks (separate pools).
- `runs/` run results (git ignores it). `runs/_archive_v1` old runs.
## 5. Commands
```
python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N
python3 -m harness.cli rescore runs/<run_dir>
python3 -m harness.cli cleanup-list <PREFIX> runs/<dir> && python3 -m harness.cli teardown runs/<dir>
python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000
python3 -m harness.pilot 2 # pilot list; `python3 -m harness.pilot 22 hard 1400` hard batch
python3 -m harness.mutation G0002 --pool eval --run-base 5000 [--keep]
python3 -c "from harness.ledger import spent; print(spent())"
```
## 6. Current state (2026-10-02)
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support;
task generator; ledger and budget guard.
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
- Pilot (G0002–G0021): 16/20 accepted, 0.16 USD per accepted task (`docs/faz1-tasarim.md` 11e).
G0020 and G0013 fixed in review and revalidated → 18/20.
- Hard batch (G0022–G0026, 2026-10-03): 5/5 accepted (G0022 after harness fixes), 0.097 USD per accepted
task, first-attempt 1/5 (`docs/faz1-tasarim.md` 11f).
- Generator: static checks, local abaplint parser check, dependency order of seed/reference, budget floor,
max_tokens 80k. Runner: G2 for CDS with parameters, G6 ignores unknown standard superclasses.
- Step D done: `harness/mutation.py`, part of generator acceptance. All 26 tasks pass (`docs/faz1-tasarim.md` 11g).
Killed mutants are kept in `<task>/faulty/`.
- Step 5 code: H stop scoring (`judge.py`), K variants (`make_k_variant`), proxy tool schema `generic_v0` (draft).
Not yet run on SAP.
- Step F: `harness/evalset.py` (88 slots G0100–G0187). Not started.
## 6a. Base model decision (2026-10-04, revised)
- **Base model: Qwen 3.8 27B** (`Qwen/Qwen3.8-27B`, Apache 2.0; weights `mlx-community/Qwen3.8-27B-4bit`,
local `~/models/Qwen3.8-27B-4bit`). Kral decision, same day, after the Devstral baseline.
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
must fix, and he does not like Devstral. Devstral is not used further.
- Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0;
end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`.
The Devstral weights were deleted. No further base model tests.
- **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook,
run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`.
Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable.
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
## 7. Next steps
1. Test one H slot (G0168) and one K slot (G0178); then run step F (`python3 -m harness.evalset run`).
2. Own-test scoring part (a): run the model's own tests against `faulty/` mutants.
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
checklist, Kral spot-checks ~10 flagged tasks.
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
5. Step 1.3b: generic ABAP MCP interface (spec `abap-mcp-arayuz.md`, later). The proxy is the first adapter.
## 8. Open items for Kral (EPOD server)
- When a write fails, run a syntax check and return its messages. Now the server says only "save failed"
(for example for `TYPE c LENGTH n` in a method signature). The model cannot see the cause and cannot repair;
this hurts eval and training. Workaround in the generator: local abaplint parser check.
- Unit test timeout: an endless loop in tested code blocks the one RFC connection for ~8 min
(mutation run 5201, G0002). Request: stop a unit test run after N seconds (DURATION SHORT = 60 s).
- Concurrency: queue calls per RFC connection, or use a connection pool.
- BDEF creation (needed for RAP tasks).