# Task: Stage 1 training (ABAP corpus) on the Mac mini ## Context Read CLAUDE.md first. This task adds stage 1 of the training plan: 1. Baseline: measure the base model on the eval tasks. 2. Stage 1: LoRA training of the base model on the ABAP corpus (continued pretraining). The model learns ABAP, CDS, RAP and DDIC syntax. 3. Measure again with the trained model. Stage 2 (SFT on teacher trajectories) is not part of this task. Input: `~/projects/abap-llm/corpus/corpus.jsonl`. Source (since 2026-10-03): SAP-samples/abap-cheat-sheets, Apache-2.0 (`corpus/out/ATTRIBUTION.md`). Real numbers from Step 1 (`train/data/report.md`): 370 records, 2,806,501 tokens (Qwen tokenizer; the chars/4 estimate is 2.64M). After version dedup and splitting: train 374 documents / 2.57M tokens, valid 20 documents / 135k tokens, 748 iterations for 2 epochs. The old numbers (1,765 documents, 1.4M tokens) are outdated. One record per line: `{"text", "objects", "types", "language_version", "package_path", "grouped", "obsolete", "tokens"}`. The `tokens` field is an estimate (chars / 4); there is also a field `source`. ## Rules - All CLAUDE.md rules apply. - Training needs the memory of the Mac mini. During training, A4H must be stopped and no other model may be loaded (also not in Ollama). Do not stop or start A4H yourself. Tell me when I must stop or start it, and wait. - Ask me before each long run (more than 30 minutes). Give the time estimate. - Do not poll long runs. Start them in the background with a log file and a macOS notification at the end, then stop and wait for my message. When I write, read only the last lines of the log. - Use the eval pool only for measurement. Never put eval tasks or their solutions into training data. - Put new code in `train/` in the harness repo. Keep changes to existing harness code small. Commit after each step. ## Step 0: Setup - Install `mlx-lm` in a separate virtual environment `train/.venv`. - Base model: the model that the harness uses for the local llm agent (Qwen 27B). Find the exact model id in the harness config. Use a 4-bit MLX version of the same model. If none exists, convert it with `mlx_lm.convert` and 4-bit quantization. Record the model id and the quantization in `train/README.md`. - Check the options of `mlx_lm.lora` with `--help`. Do not guess options. ## Step 1: Data preparation (`train/prepare.py`) - Count tokens again with the tokenizer of the base model. Replace the estimates. - Remove documents longer than `max_seq_length` (default 16384) and list them in the report. - Split by document, with a fixed seed: 95% train, 5% valid. A group is one document, so a group is never in both sets. - Write `train/data/train.jsonl` and `train/data/valid.jsonl` in the format that `mlx_lm.lora` expects for plain text (field `text`). Keep the header lines in the text. - Report: document count, token count with the real tokenizer, token distribution, removed documents. ## Step 2: Baseline (A4H must run) - Run the base model on all tasks in the eval pool that are accepted (T01, T13, T14, T15 and the accepted generated eval tasks). Use the existing harness commands. One task at a time. - Measure the base model loss on `valid.jsonl` (`mlx_lm.lora --test` without adapter, or the equivalent option). - Save the results in `runs/stage1/baseline.json`. ## Step 3: Training (A4H must be stopped) - First a short test: 20 iterations. Check memory use and time per iteration. If memory is not sufficient at 16384, tell me and suggest options (for example max_seq_length 8192, fewer LoRA layers). Do not change it yourself. - Start values (record them in `train/config.yaml`): - LoRA rank 16, all layers if memory allows - learning rate 5e-5, with warmup and cosine decay if the tool supports it - batch size 1, gradient checkpointing on - 2 epochs (iterations = train documents × 2) - valid loss every 200 iterations, save the adapter every 200 iterations - Give me the time estimate for the full run and wait for my approval. - During the run: if valid loss goes up 3 times in a row, stop and keep the adapter with the lowest valid loss. ## Step 4: Measure again (A4H must run) - Serve the base model with the adapter (for example `mlx_lm.server` with `--adapter-path`). If the harness llm agent cannot use this server, add the smallest change that lets it use an OpenAI-compatible endpoint. - Run the same eval tasks as in Step 2. - Measure the loss on `valid.jsonl` with the adapter. ## Report (`runs/stage1/report.md`, show it in the chat) - Model id, quantization, LoRA settings, iterations, training time, peak memory - Valid loss: base model and adapter - Eval results for each task: base model and adapter, score and gate results - Tool use: did the model still call the MCP tools correctly? Count tool calls with format errors, base model and adapter. - Your conclusion in 3 sentences: did stage 1 help, and is replay data necessary (it is necessary if tool use or format got worse)?