Files
abap-llm/CLAUDE.md

97 lines
5.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ABAP LLM harness — instructions for Claude Code
Talk to the user (Kral) in Turkish. Keep answers short; remove words that add no value.
When you write English text (specs, prompts, docs for the model), use ASD-STE100 Simplified Technical English.
Ask before an action that uses much cloud budget.
## 1. Project
- Goal: train an open-weight ABAP model (Apache 2.0). Base model: Apache 2.0 or MIT only.
- Role of the model: technical ABAP consultant. It writes ABAP, knows Clean ABAP and what to use how.
No SAP module knowledge; the functional side (spec or the /sapplan planner) gives it.
- Contract of the model: "ABAP task as text + generic ABAP MCP interface".
The plan format (plan-agent.md) gives the best result, but it is not mandatory.
- Training data: never Claude output. Eval tasks never go into training data.
- Read first: `docs/yol-haritasi.md` (roadmap, current position) and `docs/faz1-tasarim.md` (design).
When a decision or a result changes, update these files.
## 2. Environment (this Mac mini)
- A4H: Docker container `a4h`, HTTP `localhost:50000`, client 001, SAP_BASIS 816 SP01.
- MCP server (EPOD) runs inside ADT (Eclipse) at `127.0.0.1:3000`. Token: `.env` (`MCP_TOKEN`).
System name `A4H`, mode write.
- Ollama `127.0.0.1:11434`:
- `qwen3.8-27b-32k` (local, num_ctx 32k, ~18 GB). Slow: 20–40 min per task.
- `deepseek-v4.1-flash:cloud` (Ollama cloud, MIT): teacher and task author. ~0.05–0.08 USD per run.
- Python 3.9 (system), Node; abaplint in `node_modules`.
## 3. Rules
- Local model: one request at a time. No time limit (tool-call budget limits a run).
- Do not run local and cloud model runs at the same time: Ollama can queue them together.
- A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between
sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this.
Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs.
- Budget: `runs/ledger.jsonl` (list prices, upper bound). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`.
Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work.
- Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_`
(`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`).
Delete only objects with a run prefix. Clean up probe objects.
- The model never sees delete/teardown. The proxy (`proxy.py`) has a tool whitelist and hides other runs' objects.
- To wait for a background job, use its PID: `while kill -0 <PID> 2>/dev/null; do sleep 60; done`.
Do NOT use `pgrep -f <pattern>` in the loop: it also finds the loop's own command line, so the loop never ends.
The process name of a module run is `Python -m harness.<module>` (not `python3`); `pgrep -f "harness.pilot"`
without a loop gives the PID.
- Start long jobs with `nohup ... &` from a separate line after `cd` (a `cd ... && nohup ... &` chain runs
in a subshell; the commands after it do not see the new directory).
## 4. Layout
- `harness/mcp_client.py` MCP client (retry on 404 and on LOCK).
- `harness/adt_client.py` ADT deletion API.
- `harness/proxy.py` tool whitelist, budget, prefix filter, trajectory log.
- `harness/agents.py` oracle, null, llm (OpenAI-compatible; retries; ledger).
- `harness/runner.py` setup → agent → gates G1–G6 → hidden tests → own tests → ATC → abaplint → score → teardown.
- `harness/generator.py` task generator (cloud model writes a bundle; validation oracle = 100, null = 0; max 3 repairs).
- `harness/pilot.py` pilot list (20 tasks G0002–G0021).
- `harness/ledger.py` cost ledger and budget guard.
- `harness/cli.py` run, rescore, teardown, teardown-all, cleanup-list.
- `tasks/` hand-written tasks T01 (CLAS), T13 (FUNC), T14 (PROG + ALV), T15 (CDS). Oracle 100, null 0 for all.
- `tasks_gen/eval/`, `tasks_gen/train/` generated tasks (separate pools).
- `runs/` run results (git ignores it). `runs/_archive_v1` old runs.
## 5. Commands
```
python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N
python3 -m harness.cli rescore runs/<run_dir>
python3 -m harness.cli cleanup-list <PREFIX> runs/<dir> && python3 -m harness.cli teardown runs/<dir>
python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000
python3 -m harness.pilot 2
python3 -c "from harness.ledger import spent; print(spent())"
```
## 6. Current state (2026-10-02)
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support;
task generator; ledger and budget guard.
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
- Running: pilot (`harness.pilot`, 20 tasks) → `runs/gen/pilot.json`, log `runs/gen/pilot.log`.
A macOS notification shows when it ends.
- First generated task G0001: not accepted after 3 attempts (activation error, seed save error,
reference failed 1 hidden test). Feedback now includes the reference write/activation errors.
## 7. Next steps
1. Analyze the pilot: acceptance rate, failure causes, cost per task. Improve the generator prompt.
2. Step D: mutation check (hidden tests must fail on a broken reference).
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
checklist, Kral spot-checks ~10 flagged tasks.
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
5. Step 1.3b: generic ABAP MCP interface (spec `abap-mcp-arayuz.md`, later). The proxy is the first adapter.
## 8. Open items for Kral (server)
- Concurrency: queue calls per RFC connection, or use a connection pool.
- BDEF creation (needed for RAP tasks).