Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 17:33:55 +02:00
parent 01e99158a9
commit 65322f27ff
44 changed files with 2137 additions and 855 deletions

View File

@@ -0,0 +1,101 @@
# Task: Stage 1 training (ABAP corpus) on the Mac mini
## Context
Read CLAUDE.md first. This task adds stage 1 of the training plan:
1. Baseline: measure the base model on the eval tasks.
2. Stage 1: LoRA training of the base model on the ABAP corpus (continued
pretraining). The model learns ABAP, CDS, RAP and DDIC syntax.
3. Measure again with the trained model.
Stage 2 (SFT on teacher trajectories) is not part of this task.
Input: `~/projects/abap-llm/corpus/corpus.jsonl` (about 1,765 documents,
about 1.4M tokens, estimate). One record per line:
`{"text", "objects", "types", "language_version", "package_path",
"grouped", "obsolete", "tokens"}`. Token counts are estimates (chars / 4).
## Rules
- All CLAUDE.md rules apply.
- Training needs the memory of the Mac mini. During training, A4H must be
stopped and no other model may be loaded (also not in Ollama). Do not stop
or start A4H yourself. Tell me when I must stop or start it, and wait.
- Ask me before each long run (more than 30 minutes). Give the time estimate.
- Do not poll long runs. Start them in the background with a log file and a
macOS notification at the end, then stop and wait for my message. When I
write, read only the last lines of the log.
- Use the eval pool only for measurement. Never put eval tasks or their
solutions into training data.
- Put new code in `train/` in the harness repo. Keep changes to existing
harness code small. Commit after each step.
## Step 0: Setup
- Install `mlx-lm` in a separate virtual environment `train/.venv`.
- Base model: the model that the harness uses for the local llm agent (Qwen
27B). Find the exact model id in the harness config. Use a 4-bit MLX
version of the same model. If none exists, convert it with
`mlx_lm.convert` and 4-bit quantization. Record the model id and the
quantization in `train/README.md`.
- Check the options of `mlx_lm.lora` with `--help`. Do not guess options.
## Step 1: Data preparation (`train/prepare.py`)
- Count tokens again with the tokenizer of the base model. Replace the
estimates.
- Remove documents longer than `max_seq_length` (default 16384) and list
them in the report.
- Split by document, with a fixed seed: 95% train, 5% valid. A group is one
document, so a group is never in both sets.
- Write `train/data/train.jsonl` and `train/data/valid.jsonl` in the format
that `mlx_lm.lora` expects for plain text (field `text`). Keep the header
lines in the text.
- Report: document count, token count with the real tokenizer, token
distribution, removed documents.
## Step 2: Baseline (A4H must run)
- Run the base model on all tasks in the eval pool that are accepted
(T01, T13, T14, T15 and the accepted generated eval tasks). Use the
existing harness commands. One task at a time.
- Measure the base model loss on `valid.jsonl` (`mlx_lm.lora --test`
without adapter, or the equivalent option).
- Save the results in `runs/stage1/baseline.json`.
## Step 3: Training (A4H must be stopped)
- First a short test: 20 iterations. Check memory use and time per
iteration. If memory is not sufficient at 16384, tell me and suggest
options (for example max_seq_length 8192, fewer LoRA layers). Do not
change it yourself.
- Start values (record them in `train/config.yaml`):
- LoRA rank 16, all layers if memory allows
- learning rate 5e-5, with warmup and cosine decay if the tool supports it
- batch size 1, gradient checkpointing on
- 2 epochs (iterations = train documents × 2)
- valid loss every 200 iterations, save the adapter every 200 iterations
- Give me the time estimate for the full run and wait for my approval.
- During the run: if valid loss goes up 3 times in a row, stop and keep the
adapter with the lowest valid loss.
## Step 4: Measure again (A4H must run)
- Serve the base model with the adapter (for example `mlx_lm.server` with
`--adapter-path`). If the harness llm agent cannot use this server, add
the smallest change that lets it use an OpenAI-compatible endpoint.
- Run the same eval tasks as in Step 2.
- Measure the loss on `valid.jsonl` with the adapter.
## Report (`runs/stage1/report.md`, show it in the chat)
- Model id, quantization, LoRA settings, iterations, training time, peak
memory
- Valid loss: base model and adapter
- Eval results for each task: base model and adapter, score and gate
results
- Tool use: did the model still call the MCP tools correctly? Count tool
calls with format errors, base model and adapter.
- Your conclusion in 3 sentences: did stage 1 help, and is replay data
necessary (it is necessary if tool use or format got worse)?