serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -25,24 +25,34 @@ stay on the **Mac mini**. The series is `python3 -m harness.localqwen` (see belo
|
||||
31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors
|
||||
```
|
||||
(These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop.
|
||||
3. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
|
||||
3. **mlx-lm 0.32.0 on the MacBook** (the baseline flags need it; an older server answers `unrecognized arguments: --temp ...`). Check:
|
||||
`~/qwen-venv/bin/python -c "import mlx_lm; print(mlx_lm.__version__)"` must print `0.32.0`. If there is no such venv:
|
||||
```sh
|
||||
# with uv (recommended): brew install uv (once)
|
||||
uv venv --python 3.11 ~/qwen-venv
|
||||
uv pip install --python ~/qwen-venv/bin/python "mlx-lm==0.32.0"
|
||||
# without uv (needs Python 3.10 or newer, check python3 --version; Homebrew: brew install python@3.11):
|
||||
python3.11 -m venv ~/qwen-venv && ~/qwen-venv/bin/pip install "mlx-lm==0.32.0"
|
||||
```
|
||||
The start script uses `~/qwen-venv` first and refuses to start with an older server (clear message, nothing half started).
|
||||
4. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
|
||||
(path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh`
|
||||
(192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`.
|
||||
The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB;
|
||||
only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w <server pid>` keeps the MacBook awake as long as the server lives.
|
||||
4. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
|
||||
5. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
|
||||
```sh
|
||||
nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 &
|
||||
```
|
||||
(use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask
|
||||
"Do you want the application Python to accept incoming network connections?": **Allow**.
|
||||
5. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
|
||||
6. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
|
||||
`curl -s http://127.0.0.1:8080/v1/models` returns the model path.
|
||||
6. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
|
||||
7. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
|
||||
Both Macs must be in the same network (the mini is 192.168.178.x).
|
||||
7. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
|
||||
8. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
|
||||
`/Users/I301710/models/Qwen3.8-27B-4bit`.
|
||||
8. **Start the series on the Mac mini:**
|
||||
9. **Start the series on the Mac mini:**
|
||||
```sh
|
||||
cd ~/projects/abap-llm/harness
|
||||
python3 -m harness.localqwen --base-url http://<MacBook IP>:8080/v1 --wait 600
|
||||
|
||||
@@ -4,14 +4,14 @@
|
||||
# MacBook awake as long as the server process lives (caffeinate -w <server pid>). Usage: ~/serve_remote.sh Stop: kill $(cat ~/qwen_server.pid)
|
||||
MODEL="$HOME/models/Qwen3.8-27B-4bit"
|
||||
[ -f "$MODEL/config.json" ] || { echo "model not found: $MODEL" >&2; exit 1; }
|
||||
# the mlx_lm of the harness venv if the repo is on this Mac, else the one on PATH
|
||||
if [ -x "$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" ]; then
|
||||
SERVER="$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server"
|
||||
elif command -v mlx_lm.server >/dev/null 2>&1; then
|
||||
SERVER="$(command -v mlx_lm.server)"
|
||||
else
|
||||
echo "mlx_lm.server not found (pip install mlx-lm==0.32.0)" >&2; exit 1
|
||||
fi
|
||||
# the mlx_lm server: $MLX_SERVER, else ~/qwen-venv (mlx-lm 0.32.0, see docs/remote-model.md), else the harness venv, else PATH
|
||||
for c in "$MLX_SERVER" "$HOME/qwen-venv/bin/mlx_lm.server" "$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" "$(command -v mlx_lm.server)"; do
|
||||
if [ -n "$c" ] && [ -x "$c" ]; then SERVER="$c"; break; fi
|
||||
done
|
||||
[ -n "$SERVER" ] || { echo "mlx_lm.server not found: create ~/qwen-venv with mlx-lm 0.32.0 (docs/remote-model.md)" >&2; exit 1; }
|
||||
# the baseline flags need mlx-lm 0.32.x; an older server does not know them
|
||||
"$SERVER" --help 2>&1 | grep -q -- "--prompt-cache-bytes" || { echo "$SERVER is too old (no --prompt-cache-bytes). Install mlx-lm 0.32.0 in ~/qwen-venv (docs/remote-model.md)" >&2; exit 1; }
|
||||
echo "server: $SERVER"
|
||||
"$SERVER" \
|
||||
--model "$MODEL" \
|
||||
--host 0.0.0.0 --port 8080 \
|
||||
|
||||
Reference in New Issue
Block a user