Files
abap-llm/train/README.md

73 lines
4.8 KiB
Markdown

# Stage 1 training (train/)
Task: `docs/stage1-training-task.md`. State of the work: this file and `train/STATE.md` (later steps).
## Environment
- `train/.venv`: Python 3.11.17 (created with `uv`, Homebrew; the system Python 3.9 is not changed).
- `mlx-lm` 0.32.0, `mlx` 0.32.3.
## Base model
- Base model: `Qwen/Qwen3.8-27B` (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). It is the same base
model as the Ollama model `qwen3.8-27b-32k` that the harness used before.
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (Hugging Face), local path `~/models/Qwen3.8-27B-4bit`.
Quantization: MLX affine, 4 bit, group size 64. Size 16.1 GB. Downloaded 2026-10-03.
- Why not the Ollama weights (`qwen3.8:27b-mlx`): they are NVFP4 (modelopt) with one global scale per
layer. `mx.quantized_matmul` has no global scale, so `mlx_lm` cannot load them without a re-quantization
(a different model). Kral approved the Hugging Face download (2026-10-03).
## Serving (`train/serve.sh`)
`mlx_lm.server` at `http://127.0.0.1:8080/v1` (OpenAI-compatible). Base model without adapter; after
training the same script with `--adapter-path`.
Settings, the same as the earlier Ollama runs (runs 103, 203):
| Setting | Ollama run | mlx_lm.server |
|---|---|---|
| Thinking | on, Ollama default level `medium` | `--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}'` (template default would be `xhigh`) |
| temperature | 0.2 (sent by the harness llm agent; overrides the Modelfile value 1) | 0.2 (sent by the agent; server default `--temp 0.2`) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (Modelfile) | `--top-p 0.95 --top-k 20 --min-p 0` |
| presence / repeat penalty | 0 / 1 (neutral) | not set (neutral) |
| Output limit | none (context `num_ctx` 32768) | `--max-tokens 32768` (server default would be 512) |
| Context | 32768 | no fixed limit (memory) |
Thinking in the earlier Ollama runs: verified from run 103 (a turn with 3480 completion tokens and about
100 visible tokens).
Tool-call test (2026-10-03): one request with the harness system prompt and the MCP tool schemas returned
`sap_pull_source(objectType=INTF, objectName=ZIF_DEMO_CHECK)` in the OpenAI `tool_calls` format; the
arguments parse as JSON. The reasoning comes in a separate `reasoning` field. The chat template uses the
qwen3_coder XML tool format; `mlx_lm` parses it. First request: 86 s (prompt of 9.5k tokens).
## Eval subset
`train/subset.json` (copy: `runs/stage1/subset.json`): **11 tasks** (Kral decision 2026-10-03): one task per
category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (`train/subset_v1_25.json`) was cut because
one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128,
G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.
## Settings of the stage 1 runs (the same for baseline and after training)
| Setting | Value |
|---|---|
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | **off (`enable_thinking: false`), fixed for baseline and after training.** Sent per request as `chat_template_kwargs` by `train/baseline.py` (the server default stays thinking on). Reason: with thinking on, all 3 baseline runs ended with empty responses at the thinking limit (`runs/stage1/baseline_thinking_on.json`) |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
| Loop guard | `loop_guard` 3: the run ends when `sap_push_source` pushes the same source (object + md5) 3 times in a row; final report "Stopped: loop ...", `end_reason` "loop". Scores use the final state, so they do not change. Same after training. T01 of the first run (started before the guard) ran without it |
| Per run record | `end_reason` (report, loop, empty_response, tool_budget, time_budget, model_error, max_turns), `activation_failures`, `activation_error_messages` (unique), also in `baseline.json` |
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
| Docker (A4H) | VM memory 36 GB (`MemoryMiB` 36864), container `--memory 32g --memory-swap 32g` (2026-10-04) |
| Prompt cache of the server | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7), the
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thinking-on runs are not comparable.
## Baseline
- Runner: `python3 train/baseline.py --label baseline --run-base 20200` (one task at a time; results
`runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`). Started by `train/baseline_chain.sh`.