serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 22:16:56 +02:00
parent 6984fc7917
commit 5c07026bbb
2 changed files with 24 additions and 14 deletions

View File

@@ -25,24 +25,34 @@ stay on the **Mac mini**. The series is `python3 -m harness.localqwen` (see belo
31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors
```
(These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop.
3. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
3. **mlx-lm 0.32.0 on the MacBook** (the baseline flags need it; an older server answers `unrecognized arguments: --temp ...`). Check:
`~/qwen-venv/bin/python -c "import mlx_lm; print(mlx_lm.__version__)"` must print `0.32.0`. If there is no such venv:
```sh
# with uv (recommended): brew install uv (once)
uv venv --python 3.11 ~/qwen-venv
uv pip install --python ~/qwen-venv/bin/python "mlx-lm==0.32.0"
# without uv (needs Python 3.10 or newer, check python3 --version; Homebrew: brew install python@3.11):
python3.11 -m venv ~/qwen-venv && ~/qwen-venv/bin/pip install "mlx-lm==0.32.0"
```
The start script uses `~/qwen-venv` first and refuses to start with an older server (clear message, nothing half started).
4. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
(path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh`
(192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`.
The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB;
only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w <server pid>` keeps the MacBook awake as long as the server lives.
4. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
5. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
```sh
nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 &
```
(use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask
"Do you want the application Python to accept incoming network connections?": **Allow**.
5. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
6. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
`curl -s http://127.0.0.1:8080/v1/models` returns the model path.
6. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
7. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
Both Macs must be in the same network (the mini is 192.168.178.x).
7. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
8. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
`/Users/I301710/models/Qwen3.8-27B-4bit`.
8. **Start the series on the Mac mini:**
9. **Start the series on the Mac mini:**
```sh
cd ~/projects/abap-llm/harness
python3 -m harness.localqwen --base-url http://<MacBook IP>:8080/v1 --wait 600