Devstral Small 2: serve.sh, baseline.py without thinking args, repair rate, activation error message fix, stop rule in chain (run base 21000)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-04 08:57:12 +02:00
parent e84ea43a3d
commit 8d1db9c67c
5 changed files with 56 additions and 21 deletions

View File

@@ -1,13 +1,12 @@
#!/bin/sh
# Serve the base model (no adapter) with mlx_lm.server. Settings: train/README.md.
# Base model: Devstral Small 2 (no thinking mode, no chat-template args). Serve the base model (no adapter) with mlx_lm.server. Settings: train/README.md.
# Usage: train/serve.sh [--adapter-path PATH]
# Prompt cache limit: without it the cache grew to 7.9 GB in 15 min (2026-10-03); A4H needs 22 GB.
cd "$(dirname "$0")/.."
exec train/.venv/bin/mlx_lm.server \
--model "$HOME/models/Qwen3.8-27B-4bit" \
--model "$HOME/models/Devstral-Small-2-24B-4bit" \
--host 127.0.0.1 --port 8080 \
--temp 0.2 --top-p 0.95 --top-k 20 --min-p 0 \
--max-tokens 32768 \
--prompt-cache-size 4 --prompt-cache-bytes 6000000000 \
--chat-template-args '{"enable_thinking": true, "reasoning_effort": "medium"}' \
"$@"