Harness: runner, generator, tasks T01 T13 T14 T15, CLAUDE.md
This commit is contained in:
90
CLAUDE.md
Normal file
90
CLAUDE.md
Normal file
@@ -0,0 +1,90 @@
|
||||
# ABAP LLM harness — instructions for Claude Code
|
||||
|
||||
Talk to the user (Kral) in Turkish. Keep answers short; remove words that add no value.
|
||||
When you write English text (specs, prompts, docs for the model), use ASD-STE100 Simplified Technical English.
|
||||
Ask before an action that uses much cloud budget.
|
||||
|
||||
## 1. Project
|
||||
|
||||
- Goal: train an open-weight ABAP model (Apache 2.0). Base model: Apache 2.0 or MIT only.
|
||||
- Role of the model: technical ABAP consultant. It writes ABAP, knows Clean ABAP and what to use how.
|
||||
No SAP module knowledge; the functional side (spec or the /sapplan planner) gives it.
|
||||
- Contract of the model: "ABAP task as text + generic ABAP MCP interface".
|
||||
The plan format (plan-agent.md) gives the best result, but it is not mandatory.
|
||||
- Training data: never Claude output. Eval tasks never go into training data.
|
||||
- Read first: `docs/yol-haritasi.md` (roadmap, current position) and `docs/faz1-tasarim.md` (design).
|
||||
When a decision or a result changes, update these files.
|
||||
|
||||
## 2. Environment (this Mac mini)
|
||||
|
||||
- A4H: Docker container `a4h`, HTTP `localhost:50000`, client 001, SAP_BASIS 816 SP01.
|
||||
- MCP server (EPOD) runs inside ADT (Eclipse) at `127.0.0.1:3000`. Token: `.env` (`MCP_TOKEN`).
|
||||
System name `A4H`, mode write.
|
||||
- Ollama `127.0.0.1:11434`:
|
||||
- `qwen3.8-27b-32k` (local, num_ctx 32k, ~18 GB). Slow: 20–40 min per task.
|
||||
- `deepseek-v4.1-flash:cloud` (Ollama cloud, MIT): teacher and task author. ~0.05–0.08 USD per run.
|
||||
- Python 3.9 (system), Node; abaplint in `node_modules`.
|
||||
|
||||
## 3. Rules
|
||||
|
||||
- Local model: one request at a time. No time limit (tool-call budget limits a run).
|
||||
- Do not run local and cloud model runs at the same time: Ollama can queue them together.
|
||||
- A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between
|
||||
sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this.
|
||||
Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs.
|
||||
- Budget: `runs/ledger.jsonl` (list prices, upper bound). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`.
|
||||
Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work.
|
||||
- Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_`
|
||||
(`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`).
|
||||
Delete only objects with a run prefix. Clean up probe objects.
|
||||
- The model never sees delete/teardown. The proxy (`proxy.py`) has a tool whitelist and hides other runs' objects.
|
||||
|
||||
## 4. Layout
|
||||
|
||||
- `harness/mcp_client.py` MCP client (retry on 404 and on LOCK).
|
||||
- `harness/adt_client.py` ADT deletion API.
|
||||
- `harness/proxy.py` tool whitelist, budget, prefix filter, trajectory log.
|
||||
- `harness/agents.py` oracle, null, llm (OpenAI-compatible; retries; ledger).
|
||||
- `harness/runner.py` setup → agent → gates G1–G6 → hidden tests → own tests → ATC → abaplint → score → teardown.
|
||||
- `harness/generator.py` task generator (cloud model writes a bundle; validation oracle = 100, null = 0; max 3 repairs).
|
||||
- `harness/pilot.py` pilot list (20 tasks G0002–G0021).
|
||||
- `harness/ledger.py` cost ledger and budget guard.
|
||||
- `harness/cli.py` run, rescore, teardown, teardown-all, cleanup-list.
|
||||
- `tasks/` hand-written tasks T01 (CLAS), T13 (FUNC), T14 (PROG + ALV), T15 (CDS). Oracle 100, null 0 for all.
|
||||
- `tasks_gen/eval/`, `tasks_gen/train/` generated tasks (separate pools).
|
||||
- `runs/` run results (git ignores it). `runs/_archive_v1` old runs.
|
||||
|
||||
## 5. Commands
|
||||
|
||||
```
|
||||
python3 -m harness.cli run T01 --agent oracle|null|llm [--model NAME] --run N
|
||||
python3 -m harness.cli rescore runs/<run_dir>
|
||||
python3 -m harness.cli cleanup-list <PREFIX> runs/<dir> && python3 -m harness.cli teardown runs/<dir>
|
||||
python3 -m harness.generator --id G0100 --pool eval --object-type CLAS --category C --run-base 2000
|
||||
python3 -m harness.pilot 2
|
||||
python3 -c "from harness.ledger import spent; print(spent())"
|
||||
```
|
||||
|
||||
## 6. Current state (2026-10-02)
|
||||
|
||||
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support;
|
||||
task generator; ledger and budget guard.
|
||||
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
|
||||
- Running: pilot (`harness.pilot`, 20 tasks) → `runs/gen/pilot.json`, log `runs/gen/pilot.log`.
|
||||
A macOS notification shows when it ends.
|
||||
- First generated task G0001: not accepted after 3 attempts (activation error, seed save error,
|
||||
reference failed 1 hidden test). Feedback now includes the reference write/activation errors.
|
||||
|
||||
## 7. Next steps
|
||||
|
||||
1. Analyze the pilot: acceptance rate, failure causes, cost per task. Improve the generator prompt.
|
||||
2. Step D: mutation check (hidden tests must fail on a broken reference).
|
||||
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
|
||||
checklist, Kral spot-checks ~10 flagged tasks.
|
||||
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
|
||||
5. Step 1.3b: generic ABAP MCP interface (spec `abap-mcp-arayuz.md`, later). The proxy is the first adapter.
|
||||
|
||||
## 8. Open items for Kral (server)
|
||||
|
||||
- Concurrency: queue calls per RFC connection, or use a connection pool.
|
||||
- BDEF creation (needed for RAP tasks).
|
||||
Reference in New Issue
Block a user