Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -10,17 +10,16 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
|
||||
## Base model
|
||||
|
||||
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
|
||||
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
|
||||
(Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
|
||||
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
|
||||
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
|
||||
cannot load them without a re-quantization.
|
||||
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
|
||||
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
|
||||
- Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0,
|
||||
loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests.
|
||||
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
||||
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
||||
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
|
||||
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
|
||||
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
|
||||
`train/serve.sh` serves Qwen (commit a2ba9e7).
|
||||
|
||||
## Serving (`train/serve.sh`)
|
||||
|
||||
@@ -29,17 +28,14 @@ training the same script with `--adapter-path`.
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent |
|
||||
| Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request |
|
||||
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
|
||||
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
|
||||
| presence / repeat penalty | not set (neutral) |
|
||||
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
|
||||
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
||||
|
||||
Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`);
|
||||
`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors.
|
||||
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
|
||||
(G0128, G0174 ended with `report`).
|
||||
Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed).
|
||||
|
||||
## Eval subset
|
||||
|
||||
@@ -52,8 +48,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) |
|
||||
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) |
|
||||
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
||||
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
||||
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
||||
@@ -68,12 +64,11 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
|
||||
|
||||
## Baseline
|
||||
|
||||
- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by
|
||||
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
|
||||
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
|
||||
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
|
||||
- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
|
||||
(the model changes the source, but the changes do not remove the cause).
|
||||
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
||||
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).
|
||||
- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.
|
||||
- Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off).
|
||||
Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`.
|
||||
Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1.
|
||||
Repair rate 39/54 = 0.72.
|
||||
- Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84.
|
||||
Results in `runs/archive/devstral/`.
|
||||
- Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
||||
- The stage 1 training test (step 3) uses Qwen (`config_test.yaml`).
|
||||
|
||||
Reference in New Issue
Block a user