Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
83
docs/stage1-step-prompts.md
Normal file
83
docs/stage1-step-prompts.md
Normal file
@@ -0,0 +1,83 @@
|
||||
# Stage 1: step prompts
|
||||
|
||||
Preparation (one time): copy `stage1-training-task.md` to
|
||||
`~/projects/abap-llm/harness/docs/stage1-training-task.md`.
|
||||
|
||||
Use Sonnet. Before each prompt: `/clear`. Each prompt is one session.
|
||||
|
||||
---
|
||||
|
||||
## A — Setup and data (A4H: no change needed)
|
||||
|
||||
```
|
||||
Read CLAUDE.md and docs/stage1-training-task.md. Do only Step 0 and Step 1.
|
||||
Create train/STATE.md: write what you did, the model id, the paths, the
|
||||
report of Step 1, and the next step. Commit. Then stop.
|
||||
```
|
||||
|
||||
## B1 — Baseline, start (A4H must run)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Do Step 2.
|
||||
Start the eval runs in the background with a log file and a macOS
|
||||
notification at the end. Measure the valid loss of the base model in the
|
||||
same background job, after the eval runs. Update train/STATE.md with the
|
||||
log path and the job command. Then stop. Do not poll.
|
||||
```
|
||||
|
||||
## B2 — Baseline, results (after the notification)
|
||||
|
||||
```
|
||||
Read CLAUDE.md and train/STATE.md. Read only the last 50 lines of the log.
|
||||
Write runs/stage1/baseline.json. Update train/STATE.md with the results.
|
||||
Commit. Show me a short summary. Then stop.
|
||||
```
|
||||
|
||||
## C — Training test (stop A4H first)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. A4H is
|
||||
stopped. Check that no other model is loaded. Do only the short test of
|
||||
Step 3 (20 iterations). Report peak memory, time per iteration, and the
|
||||
time estimate for the full run. If memory is not sufficient, give me
|
||||
options and do not change the settings. Update train/STATE.md. Then stop.
|
||||
```
|
||||
|
||||
## D1 — Training, start (after my approval)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Start the
|
||||
full training run of Step 3 in the background with a log file and a macOS
|
||||
notification at the end. Use the settings from train/config.yaml. Update
|
||||
train/STATE.md with the log path and the adapter path. Then stop. Do not
|
||||
poll.
|
||||
```
|
||||
|
||||
## D2 — Training, results (after the notification)
|
||||
|
||||
```
|
||||
Read CLAUDE.md and train/STATE.md. Read only the valid loss lines and the
|
||||
last 30 lines of the log. Select the adapter with the lowest valid loss.
|
||||
Update train/STATE.md. Commit. Show me the valid loss curve as a short
|
||||
table. Then stop.
|
||||
```
|
||||
|
||||
## E1 — Measure again, start (start A4H first)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. A4H runs.
|
||||
Do Step 4: serve the base model with the selected adapter, connect the
|
||||
harness llm agent to it (smallest change), and test it with 1 task. Then
|
||||
start the same eval runs as in Step 2 and the valid loss measurement in the
|
||||
background with a log file and a macOS notification at the end. Update
|
||||
train/STATE.md. Commit. Then stop. Do not poll.
|
||||
```
|
||||
|
||||
## E2 — Report (after the notification)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Read only
|
||||
the result files and the last 50 lines of the log. Write
|
||||
runs/stage1/report.md as described in the task. Update train/STATE.md.
|
||||
Commit. Show me the report. Then stop.
|
||||
```
|
||||
101
docs/stage1-training-task.md
Normal file
101
docs/stage1-training-task.md
Normal file
@@ -0,0 +1,101 @@
|
||||
# Task: Stage 1 training (ABAP corpus) on the Mac mini
|
||||
|
||||
## Context
|
||||
|
||||
Read CLAUDE.md first. This task adds stage 1 of the training plan:
|
||||
|
||||
1. Baseline: measure the base model on the eval tasks.
|
||||
2. Stage 1: LoRA training of the base model on the ABAP corpus (continued
|
||||
pretraining). The model learns ABAP, CDS, RAP and DDIC syntax.
|
||||
3. Measure again with the trained model.
|
||||
|
||||
Stage 2 (SFT on teacher trajectories) is not part of this task.
|
||||
|
||||
Input: `~/projects/abap-llm/corpus/corpus.jsonl` (about 1,765 documents,
|
||||
about 1.4M tokens, estimate). One record per line:
|
||||
`{"text", "objects", "types", "language_version", "package_path",
|
||||
"grouped", "obsolete", "tokens"}`. Token counts are estimates (chars / 4).
|
||||
|
||||
## Rules
|
||||
|
||||
- All CLAUDE.md rules apply.
|
||||
- Training needs the memory of the Mac mini. During training, A4H must be
|
||||
stopped and no other model may be loaded (also not in Ollama). Do not stop
|
||||
or start A4H yourself. Tell me when I must stop or start it, and wait.
|
||||
- Ask me before each long run (more than 30 minutes). Give the time estimate.
|
||||
- Do not poll long runs. Start them in the background with a log file and a
|
||||
macOS notification at the end, then stop and wait for my message. When I
|
||||
write, read only the last lines of the log.
|
||||
- Use the eval pool only for measurement. Never put eval tasks or their
|
||||
solutions into training data.
|
||||
- Put new code in `train/` in the harness repo. Keep changes to existing
|
||||
harness code small. Commit after each step.
|
||||
|
||||
## Step 0: Setup
|
||||
|
||||
- Install `mlx-lm` in a separate virtual environment `train/.venv`.
|
||||
- Base model: the model that the harness uses for the local llm agent (Qwen
|
||||
27B). Find the exact model id in the harness config. Use a 4-bit MLX
|
||||
version of the same model. If none exists, convert it with
|
||||
`mlx_lm.convert` and 4-bit quantization. Record the model id and the
|
||||
quantization in `train/README.md`.
|
||||
- Check the options of `mlx_lm.lora` with `--help`. Do not guess options.
|
||||
|
||||
## Step 1: Data preparation (`train/prepare.py`)
|
||||
|
||||
- Count tokens again with the tokenizer of the base model. Replace the
|
||||
estimates.
|
||||
- Remove documents longer than `max_seq_length` (default 16384) and list
|
||||
them in the report.
|
||||
- Split by document, with a fixed seed: 95% train, 5% valid. A group is one
|
||||
document, so a group is never in both sets.
|
||||
- Write `train/data/train.jsonl` and `train/data/valid.jsonl` in the format
|
||||
that `mlx_lm.lora` expects for plain text (field `text`). Keep the header
|
||||
lines in the text.
|
||||
- Report: document count, token count with the real tokenizer, token
|
||||
distribution, removed documents.
|
||||
|
||||
## Step 2: Baseline (A4H must run)
|
||||
|
||||
- Run the base model on all tasks in the eval pool that are accepted
|
||||
(T01, T13, T14, T15 and the accepted generated eval tasks). Use the
|
||||
existing harness commands. One task at a time.
|
||||
- Measure the base model loss on `valid.jsonl` (`mlx_lm.lora --test`
|
||||
without adapter, or the equivalent option).
|
||||
- Save the results in `runs/stage1/baseline.json`.
|
||||
|
||||
## Step 3: Training (A4H must be stopped)
|
||||
|
||||
- First a short test: 20 iterations. Check memory use and time per
|
||||
iteration. If memory is not sufficient at 16384, tell me and suggest
|
||||
options (for example max_seq_length 8192, fewer LoRA layers). Do not
|
||||
change it yourself.
|
||||
- Start values (record them in `train/config.yaml`):
|
||||
- LoRA rank 16, all layers if memory allows
|
||||
- learning rate 5e-5, with warmup and cosine decay if the tool supports it
|
||||
- batch size 1, gradient checkpointing on
|
||||
- 2 epochs (iterations = train documents × 2)
|
||||
- valid loss every 200 iterations, save the adapter every 200 iterations
|
||||
- Give me the time estimate for the full run and wait for my approval.
|
||||
- During the run: if valid loss goes up 3 times in a row, stop and keep the
|
||||
adapter with the lowest valid loss.
|
||||
|
||||
## Step 4: Measure again (A4H must run)
|
||||
|
||||
- Serve the base model with the adapter (for example `mlx_lm.server` with
|
||||
`--adapter-path`). If the harness llm agent cannot use this server, add
|
||||
the smallest change that lets it use an OpenAI-compatible endpoint.
|
||||
- Run the same eval tasks as in Step 2.
|
||||
- Measure the loss on `valid.jsonl` with the adapter.
|
||||
|
||||
## Report (`runs/stage1/report.md`, show it in the chat)
|
||||
|
||||
- Model id, quantization, LoRA settings, iterations, training time, peak
|
||||
memory
|
||||
- Valid loss: base model and adapter
|
||||
- Eval results for each task: base model and adapter, score and gate
|
||||
results
|
||||
- Tool use: did the model still call the MCP tools correctly? Count tool
|
||||
calls with format errors, base model and adapter.
|
||||
- Your conclusion in 3 sentences: did stage 1 help, and is replay data
|
||||
necessary (it is necessary if tool use or format got worse)?
|
||||
Reference in New Issue
Block a user