Files
abap-llm/train/README.md

75 lines
5.1 KiB
Markdown

# Stage 1 training (train/)
Task: `docs/stage1-training-task.md`. State of the work: this file and `train/STATE.md` (later steps).
## Environment
- `train/.venv`: Python 3.11.17 (created with `uv`, Homebrew; the system Python 3.9 is not changed).
- `mlx-lm` 0.32.0, `mlx` 0.32.3.
## Base model
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
(Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
cannot load them without a re-quantization.
- Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0,
loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests.
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
`train/serve.sh` serves Qwen (commit a2ba9e7).
## Serving (`train/serve.sh`)
`mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after
training the same script with `--adapter-path`.
| Setting | Value |
|---|---|
| Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request |
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
| presence / repeat penalty | not set (neutral) |
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed).
## Eval subset
`train/subset.json` (copy: `runs/stage1/subset.json`): **11 tasks** (Kral decision 2026-10-03): one task per
category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (`train/subset_v1_25.json`) was cut because
one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128,
G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.
## Settings of the stage 1 runs (the same for baseline and after training)
| Setting | Value |
|---|---|
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
| Loop guard | `loop_guard` 3: the run ends when `sap_push_source` pushes the same source (object + md5) 3 times in a row, or when any tool is called 3 times in a row with the same arguments and the same result (read loop, e.g. G0105 pulled the same include 18 times); final report "Stopped: loop ...", `end_reason` "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it |
| Per run record | `end_reason` (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), `activation_failures`, `activation_error_messages` (unique), also in `baseline.json` |
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
| Docker (A4H) | VM memory 36 GB (`MemoryMiB` 36864), container `--memory 32g --memory-swap 32g` (2026-10-04) |
| Prompt cache of the server | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7), the
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thinking-on runs are not comparable.
## Baseline
- Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off).
Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`.
Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1.
Repair rate 39/54 = 0.72.
- Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84.
Results in `runs/archive/devstral/`.
- Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
- The stage 1 training test (step 3) uses Qwen (`config_test.yaml`).