Generator: pilot analysis and fixes; G0002-G0021 tasks; handover notes merged

- Static bundle checks (name length, seed type, reserved words, contract test classes,
  testclasses_file), local abaplint parser check before SAP, max_tokens, robust JSON parse
- Runner: G2 for CDS views with parameters, contract_check detail in report
- G0020 and G0013 fixed by review and revalidated (oracle 100, null 0)
- Pilot results in docs/faz1-tasarim.md 11e; handover notes merged into CLAUDE.md and docs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-03 05:04:23 +02:00
parent 15eb9f1bb2
commit 5de5e2851c
162 changed files with 9943 additions and 23 deletions

View File

@@ -1,8 +1,15 @@
# ABAP LLM harness — instructions for Claude Code
Talk to the user (Kral) in Turkish. Keep answers short; remove words that add no value.
When you write English text (specs, prompts, docs for the model), use ASD-STE100 Simplified Technical English.
Ask before an action that uses much cloud budget.
## 0. Working with Kral
- Talk in Turkish. Short answers. Remove words and sentences that add no value.
- English text (specs, prompts, docs for the model): ASD-STE100 Simplified Technical English.
- Ask before an action that costs much (cloud budget) or that changes the A4H system outside harness runs.
- For a series of updates to one document, deliver the result directly as a Markdown file; do not ask each time.
- Record each decision with its date in `docs/yol-haritasi.md` or `docs/faz1-tasarim.md`.
- Commit the current changes before a new generation or test batch.
- Kral has no time for a full review. Eval review = automatic gates + empirical filter (2–3 models) +
Claude review with a checklist; Kral reviews only flagged tasks + one per category (~10 tasks).
## 1. Project
@@ -12,6 +19,12 @@ Ask before an action that uses much cloud budget.
- Contract of the model: "ABAP task as text + generic ABAP MCP interface".
The plan format (plan-agent.md) gives the best result, but it is not mandatory.
- Training data: never Claude output. Eval tasks never go into training data.
- Input: the plan format gives the best result but is not mandatory. The old playbook format is not used.
- Generic ABAP MCP interface (core: search_object, read_source, object_structure, where_used, create_object,
write_source, activate, syntax_check, run_unit_tests, run_atc). Small tool name/schema variation in training.
- Category K: incomplete or free-text input + other tool schema.
- Object type mix: CLAS/INTF 35 %, FUNC 15 %, PROG 10 %, DDIC 10 %, CDS 25 %, MSAG + exception 5 %.
- Teacher and task author: DeepSeek V4.1 Flash (Ollama cloud). Budget calendar in `docs/yol-haritasi.md`.
- Read first: `docs/yol-haritasi.md` (roadmap, current position) and `docs/faz1-tasarim.md` (design).
When a decision or a result changes, update these files.
@@ -76,21 +89,27 @@ python3 -c "from harness.ledger import spent; print(spent())"
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support;
task generator; ledger and budget guard.
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
- Running: pilot (`harness.pilot`, 20 tasks) → `runs/gen/pilot.json`, log `runs/gen/pilot.log`.
A macOS notification shows when it ends.
- First generated task G0001: not accepted after 3 attempts (activation error, seed save error,
reference failed 1 hidden test). Feedback now includes the reference write/activation errors.
- Pilot (G0002–G0021) done: 16/20 accepted, 0.16 USD per accepted task (`runs/gen/pilot.json`,
analysis in `docs/faz1-tasarim.md` 11e). Generator improved after the pilot: static checks
(name length, seed type, reserved words, contract/test classes), local abaplint parser check before SAP,
max_tokens 100k, G2 detail in the repair feedback. Runner: G2 works for CDS views with parameters.
The improved generator is not yet tested with new tasks.
- G0020 (rejected) passes with oracle 100 after one fix (remove parameter `default`). G0013 (accepted) has a test
class in its contract: fix in review.
## 7. Next steps
1. Analyze the pilot: acceptance rate, failure causes, cost per task. Improve the generator prompt.
1. Test the improved generator with a small batch (~5 tasks, ~0.7 USD); compare repairs per task with the pilot (2.75 calls).
2. Step D: mutation check (hidden tests must fail on a broken reference).
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
checklist, Kral spot-checks ~10 flagged tasks.
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
5. Step 1.3b: generic ABAP MCP interface (spec `abap-mcp-arayuz.md`, later). The proxy is the first adapter.
## 8. Open items for Kral (server)
## 8. Open items for Kral (EPOD server)
- When a write fails, run a syntax check and return its messages. Now the server says only "save failed"
(for example for `TYPE c LENGTH n` in a method signature). The model cannot see the cause and cannot repair;
this hurts eval and training. Workaround in the generator: local abaplint parser check.
- Concurrency: queue calls per RFC connection, or use a connection pool.
- BDEF creation (needed for RAP tasks).