Baseline shortened: 11-task subset (1 per category + T01), max_tokens 16384, settings in README; baseline_chain.sh
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -43,12 +43,27 @@ qwen3_coder XML tool format; `mlx_lm` parses it. First request: 86 s (prompt of
|
||||
|
||||
## Eval subset
|
||||
|
||||
`train/subset.json` (copy: `runs/stage1/subset.json`): 25 tasks, balanced over the categories
|
||||
(A, B, C, E, F: 3 each; D, G, H, I, K: 2 each), with T01. Use the same list before and after training.
|
||||
`train/subset.json` (copy: `runs/stage1/subset.json`): **11 tasks** (Kral decision 2026-10-03): one task per
|
||||
category (A B C D E F G H I K) plus T01. The first subset of 25 tasks (`train/subset_v1_25.json`) was cut because
|
||||
one task needed 2-3 hours with the MLX 4-bit model (about 12 tokens/s). Tasks: T01, G0105, G0017, G0125, G0128,
|
||||
G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after training.
|
||||
|
||||
## Settings of the stage 1 runs (the same for baseline and after training)
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | on, `reasoning_effort` medium |
|
||||
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
||||
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
||||
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
||||
| Empty turn | retried (2 times), then the run stops ("Stopped: empty model response") |
|
||||
| Prompt cache of the server | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
||||
|
||||
The settings are also written into `runs/stage1/baseline.json` (`settings`). The old Ollama run (41.7) and the
|
||||
T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) are not comparable.
|
||||
|
||||
## Baseline
|
||||
|
||||
- Runner: `python3 train/baseline.py --label baseline` (one task at a time; results
|
||||
`runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`).
|
||||
- The earlier T01 score 41.7 (run 103) used the Ollama NVFP4 weights. The baseline of 2026-10-03 with the
|
||||
MLX 4-bit weights is the new reference for stage 1.
|
||||
- Runner: `python3 train/baseline.py --label baseline --run-base 20200` (one task at a time; results
|
||||
`runs/stage1/baseline.json`, run directories `runs/stage1/baseline/`). Started by `train/baseline_chain.sh`.
|
||||
|
||||
Reference in New Issue
Block a user