Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 17:33:55 +02:00
parent 01e99158a9
commit 65322f27ff
44 changed files with 2137 additions and 855 deletions

View File

@@ -0,0 +1,83 @@
# Stage 1: step prompts
Preparation (one time): copy `stage1-training-task.md` to
`~/projects/abap-llm/harness/docs/stage1-training-task.md`.
Use Sonnet. Before each prompt: `/clear`. Each prompt is one session.
---
## A — Setup and data (A4H: no change needed)
```
Read CLAUDE.md and docs/stage1-training-task.md. Do only Step 0 and Step 1.
Create train/STATE.md: write what you did, the model id, the paths, the
report of Step 1, and the next step. Commit. Then stop.
```
## B1 — Baseline, start (A4H must run)
```
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Do Step 2.
Start the eval runs in the background with a log file and a macOS
notification at the end. Measure the valid loss of the base model in the
same background job, after the eval runs. Update train/STATE.md with the
log path and the job command. Then stop. Do not poll.
```
## B2 — Baseline, results (after the notification)
```
Read CLAUDE.md and train/STATE.md. Read only the last 50 lines of the log.
Write runs/stage1/baseline.json. Update train/STATE.md with the results.
Commit. Show me a short summary. Then stop.
```
## C — Training test (stop A4H first)
```
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. A4H is
stopped. Check that no other model is loaded. Do only the short test of
Step 3 (20 iterations). Report peak memory, time per iteration, and the
time estimate for the full run. If memory is not sufficient, give me
options and do not change the settings. Update train/STATE.md. Then stop.
```
## D1 — Training, start (after my approval)
```
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Start the
full training run of Step 3 in the background with a log file and a macOS
notification at the end. Use the settings from train/config.yaml. Update
train/STATE.md with the log path and the adapter path. Then stop. Do not
poll.
```
## D2 — Training, results (after the notification)
```
Read CLAUDE.md and train/STATE.md. Read only the valid loss lines and the
last 30 lines of the log. Select the adapter with the lowest valid loss.
Update train/STATE.md. Commit. Show me the valid loss curve as a short
table. Then stop.
```
## E1 — Measure again, start (start A4H first)
```
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. A4H runs.
Do Step 4: serve the base model with the selected adapter, connect the
harness llm agent to it (smallest change), and test it with 1 task. Then
start the same eval runs as in Step 2 and the valid loss measurement in the
background with a log file and a macOS notification at the end. Update
train/STATE.md. Commit. Then stop. Do not poll.
```
## E2 — Report (after the notification)
```
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Read only
the result files and the last 50 lines of the log. Write
runs/stage1/report.md as described in the task. Update train/STATE.md.
Commit. Show me the report. Then stop.
```

View File

@@ -0,0 +1,101 @@
# Task: Stage 1 training (ABAP corpus) on the Mac mini
## Context
Read CLAUDE.md first. This task adds stage 1 of the training plan:
1. Baseline: measure the base model on the eval tasks.
2. Stage 1: LoRA training of the base model on the ABAP corpus (continued
pretraining). The model learns ABAP, CDS, RAP and DDIC syntax.
3. Measure again with the trained model.
Stage 2 (SFT on teacher trajectories) is not part of this task.
Input: `~/projects/abap-llm/corpus/corpus.jsonl` (about 1,765 documents,
about 1.4M tokens, estimate). One record per line:
`{"text", "objects", "types", "language_version", "package_path",
"grouped", "obsolete", "tokens"}`. Token counts are estimates (chars / 4).
## Rules
- All CLAUDE.md rules apply.
- Training needs the memory of the Mac mini. During training, A4H must be
stopped and no other model may be loaded (also not in Ollama). Do not stop
or start A4H yourself. Tell me when I must stop or start it, and wait.
- Ask me before each long run (more than 30 minutes). Give the time estimate.
- Do not poll long runs. Start them in the background with a log file and a
macOS notification at the end, then stop and wait for my message. When I
write, read only the last lines of the log.
- Use the eval pool only for measurement. Never put eval tasks or their
solutions into training data.
- Put new code in `train/` in the harness repo. Keep changes to existing
harness code small. Commit after each step.
## Step 0: Setup
- Install `mlx-lm` in a separate virtual environment `train/.venv`.
- Base model: the model that the harness uses for the local llm agent (Qwen
27B). Find the exact model id in the harness config. Use a 4-bit MLX
version of the same model. If none exists, convert it with
`mlx_lm.convert` and 4-bit quantization. Record the model id and the
quantization in `train/README.md`.
- Check the options of `mlx_lm.lora` with `--help`. Do not guess options.
## Step 1: Data preparation (`train/prepare.py`)
- Count tokens again with the tokenizer of the base model. Replace the
estimates.
- Remove documents longer than `max_seq_length` (default 16384) and list
them in the report.
- Split by document, with a fixed seed: 95% train, 5% valid. A group is one
document, so a group is never in both sets.
- Write `train/data/train.jsonl` and `train/data/valid.jsonl` in the format
that `mlx_lm.lora` expects for plain text (field `text`). Keep the header
lines in the text.
- Report: document count, token count with the real tokenizer, token
distribution, removed documents.
## Step 2: Baseline (A4H must run)
- Run the base model on all tasks in the eval pool that are accepted
(T01, T13, T14, T15 and the accepted generated eval tasks). Use the
existing harness commands. One task at a time.
- Measure the base model loss on `valid.jsonl` (`mlx_lm.lora --test`
without adapter, or the equivalent option).
- Save the results in `runs/stage1/baseline.json`.
## Step 3: Training (A4H must be stopped)
- First a short test: 20 iterations. Check memory use and time per
iteration. If memory is not sufficient at 16384, tell me and suggest
options (for example max_seq_length 8192, fewer LoRA layers). Do not
change it yourself.
- Start values (record them in `train/config.yaml`):
- LoRA rank 16, all layers if memory allows
- learning rate 5e-5, with warmup and cosine decay if the tool supports it
- batch size 1, gradient checkpointing on
- 2 epochs (iterations = train documents × 2)
- valid loss every 200 iterations, save the adapter every 200 iterations
- Give me the time estimate for the full run and wait for my approval.
- During the run: if valid loss goes up 3 times in a row, stop and keep the
adapter with the lowest valid loss.
## Step 4: Measure again (A4H must run)
- Serve the base model with the adapter (for example `mlx_lm.server` with
`--adapter-path`). If the harness llm agent cannot use this server, add
the smallest change that lets it use an OpenAI-compatible endpoint.
- Run the same eval tasks as in Step 2.
- Measure the loss on `valid.jsonl` with the adapter.
## Report (`runs/stage1/report.md`, show it in the chat)
- Model id, quantization, LoRA settings, iterations, training time, peak
memory
- Valid loss: base model and adapter
- Eval results for each task: base model and adapter, score and gate
results
- Tool use: did the model still call the MCP tools correctly? Count tool
calls with format errors, base model and adapter.
- Your conclusion in 3 sentences: did stage 1 help, and is replay data
necessary (it is necessary if tool use or format got worse)?