Generator: pilot analysis and fixes; G0002-G0021 tasks; handover notes merged
- Static bundle checks (name length, seed type, reserved words, contract test classes, testclasses_file), local abaplint parser check before SAP, max_tokens, robust JSON parse - Runner: G2 for CDS views with parameters, contract_check detail in report - G0020 and G0013 fixed by review and revalidated (oracle 100, null 0) - Pilot results in docs/faz1-tasarim.md 11e; handover notes merged into CLAUDE.md and docs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
37
CLAUDE.md
37
CLAUDE.md
@@ -1,8 +1,15 @@
|
||||
# ABAP LLM harness — instructions for Claude Code
|
||||
|
||||
Talk to the user (Kral) in Turkish. Keep answers short; remove words that add no value.
|
||||
When you write English text (specs, prompts, docs for the model), use ASD-STE100 Simplified Technical English.
|
||||
Ask before an action that uses much cloud budget.
|
||||
## 0. Working with Kral
|
||||
|
||||
- Talk in Turkish. Short answers. Remove words and sentences that add no value.
|
||||
- English text (specs, prompts, docs for the model): ASD-STE100 Simplified Technical English.
|
||||
- Ask before an action that costs much (cloud budget) or that changes the A4H system outside harness runs.
|
||||
- For a series of updates to one document, deliver the result directly as a Markdown file; do not ask each time.
|
||||
- Record each decision with its date in `docs/yol-haritasi.md` or `docs/faz1-tasarim.md`.
|
||||
- Commit the current changes before a new generation or test batch.
|
||||
- Kral has no time for a full review. Eval review = automatic gates + empirical filter (2–3 models) +
|
||||
Claude review with a checklist; Kral reviews only flagged tasks + one per category (~10 tasks).
|
||||
|
||||
## 1. Project
|
||||
|
||||
@@ -12,6 +19,12 @@ Ask before an action that uses much cloud budget.
|
||||
- Contract of the model: "ABAP task as text + generic ABAP MCP interface".
|
||||
The plan format (plan-agent.md) gives the best result, but it is not mandatory.
|
||||
- Training data: never Claude output. Eval tasks never go into training data.
|
||||
- Input: the plan format gives the best result but is not mandatory. The old playbook format is not used.
|
||||
- Generic ABAP MCP interface (core: search_object, read_source, object_structure, where_used, create_object,
|
||||
write_source, activate, syntax_check, run_unit_tests, run_atc). Small tool name/schema variation in training.
|
||||
- Category K: incomplete or free-text input + other tool schema.
|
||||
- Object type mix: CLAS/INTF 35 %, FUNC 15 %, PROG 10 %, DDIC 10 %, CDS 25 %, MSAG + exception 5 %.
|
||||
- Teacher and task author: DeepSeek V4.1 Flash (Ollama cloud). Budget calendar in `docs/yol-haritasi.md`.
|
||||
- Read first: `docs/yol-haritasi.md` (roadmap, current position) and `docs/faz1-tasarim.md` (design).
|
||||
When a decision or a result changes, update these files.
|
||||
|
||||
@@ -76,21 +89,27 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
||||
- Done: harness on Mac mini; test include creation in the server; teardown; FUNC/PROG/DDLS support;
|
||||
task generator; ledger and budget guard.
|
||||
- Results: T01 oracle 100, Qwen 27B 41.7 (old run 103), DeepSeek 98.5. T13 DeepSeek 85, T14 DeepSeek 100.
|
||||
- Running: pilot (`harness.pilot`, 20 tasks) → `runs/gen/pilot.json`, log `runs/gen/pilot.log`.
|
||||
A macOS notification shows when it ends.
|
||||
- First generated task G0001: not accepted after 3 attempts (activation error, seed save error,
|
||||
reference failed 1 hidden test). Feedback now includes the reference write/activation errors.
|
||||
- Pilot (G0002–G0021) done: 16/20 accepted, 0.16 USD per accepted task (`runs/gen/pilot.json`,
|
||||
analysis in `docs/faz1-tasarim.md` 11e). Generator improved after the pilot: static checks
|
||||
(name length, seed type, reserved words, contract/test classes), local abaplint parser check before SAP,
|
||||
max_tokens 100k, G2 detail in the repair feedback. Runner: G2 works for CDS views with parameters.
|
||||
The improved generator is not yet tested with new tasks.
|
||||
- G0020 (rejected) passes with oracle 100 after one fix (remove parameter `default`). G0013 (accepted) has a test
|
||||
class in its contract: fix in review.
|
||||
|
||||
## 7. Next steps
|
||||
|
||||
1. Analyze the pilot: acceptance rate, failure causes, cost per task. Improve the generator prompt.
|
||||
1. Test the improved generator with a small batch (~5 tasks, ~0.7 USD); compare repairs per task with the pilot (2.75 calls).
|
||||
2. Step D: mutation check (hidden tests must fail on a broken reference).
|
||||
3. Step F: eval set of 110 tasks: ~150 candidates, empirical filter (2–3 models), Claude review with a
|
||||
checklist, Kral spot-checks ~10 flagged tasks.
|
||||
4. Scoring for stop tasks (category H) and category K (free-text input, other tool schema).
|
||||
5. Step 1.3b: generic ABAP MCP interface (spec `abap-mcp-arayuz.md`, later). The proxy is the first adapter.
|
||||
|
||||
## 8. Open items for Kral (server)
|
||||
## 8. Open items for Kral (EPOD server)
|
||||
|
||||
- When a write fails, run a syntax check and return its messages. Now the server says only "save failed"
|
||||
(for example for `TYPE c LENGTH n` in a method signature). The model cannot see the cause and cannot repair;
|
||||
this hurts eval and training. Workaround in the generator: local abaplint parser check.
|
||||
- Concurrency: queue calls per RFC connection, or use a connection pool.
|
||||
- BDEF creation (needed for RAP tasks).
|
||||
|
||||
Reference in New Issue
Block a user