serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -25,24 +25,34 @@ stay on the **Mac mini**. The series is `python3 -m harness.localqwen` (see belo
|
|||||||
31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors
|
31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors
|
||||||
```
|
```
|
||||||
(These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop.
|
(These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop.
|
||||||
3. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
|
3. **mlx-lm 0.32.0 on the MacBook** (the baseline flags need it; an older server answers `unrecognized arguments: --temp ...`). Check:
|
||||||
|
`~/qwen-venv/bin/python -c "import mlx_lm; print(mlx_lm.__version__)"` must print `0.32.0`. If there is no such venv:
|
||||||
|
```sh
|
||||||
|
# with uv (recommended): brew install uv (once)
|
||||||
|
uv venv --python 3.11 ~/qwen-venv
|
||||||
|
uv pip install --python ~/qwen-venv/bin/python "mlx-lm==0.32.0"
|
||||||
|
# without uv (needs Python 3.10 or newer, check python3 --version; Homebrew: brew install python@3.11):
|
||||||
|
python3.11 -m venv ~/qwen-venv && ~/qwen-venv/bin/pip install "mlx-lm==0.32.0"
|
||||||
|
```
|
||||||
|
The start script uses `~/qwen-venv` first and refuses to start with an older server (clear message, nothing half started).
|
||||||
|
4. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
|
||||||
(path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh`
|
(path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh`
|
||||||
(192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`.
|
(192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`.
|
||||||
The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB;
|
The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB;
|
||||||
only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w <server pid>` keeps the MacBook awake as long as the server lives.
|
only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w <server pid>` keeps the MacBook awake as long as the server lives.
|
||||||
4. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
|
5. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
|
||||||
```sh
|
```sh
|
||||||
nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 &
|
nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 &
|
||||||
```
|
```
|
||||||
(use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask
|
(use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask
|
||||||
"Do you want the application Python to accept incoming network connections?": **Allow**.
|
"Do you want the application Python to accept incoming network connections?": **Allow**.
|
||||||
5. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
|
6. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
|
||||||
`curl -s http://127.0.0.1:8080/v1/models` returns the model path.
|
`curl -s http://127.0.0.1:8080/v1/models` returns the model path.
|
||||||
6. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
|
7. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
|
||||||
Both Macs must be in the same network (the mini is 192.168.178.x).
|
Both Macs must be in the same network (the mini is 192.168.178.x).
|
||||||
7. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
|
8. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
|
||||||
`/Users/I301710/models/Qwen3.8-27B-4bit`.
|
`/Users/I301710/models/Qwen3.8-27B-4bit`.
|
||||||
8. **Start the series on the Mac mini:**
|
9. **Start the series on the Mac mini:**
|
||||||
```sh
|
```sh
|
||||||
cd ~/projects/abap-llm/harness
|
cd ~/projects/abap-llm/harness
|
||||||
python3 -m harness.localqwen --base-url http://<MacBook IP>:8080/v1 --wait 600
|
python3 -m harness.localqwen --base-url http://<MacBook IP>:8080/v1 --wait 600
|
||||||
|
|||||||
@@ -4,14 +4,14 @@
|
|||||||
# MacBook awake as long as the server process lives (caffeinate -w <server pid>). Usage: ~/serve_remote.sh Stop: kill $(cat ~/qwen_server.pid)
|
# MacBook awake as long as the server process lives (caffeinate -w <server pid>). Usage: ~/serve_remote.sh Stop: kill $(cat ~/qwen_server.pid)
|
||||||
MODEL="$HOME/models/Qwen3.8-27B-4bit"
|
MODEL="$HOME/models/Qwen3.8-27B-4bit"
|
||||||
[ -f "$MODEL/config.json" ] || { echo "model not found: $MODEL" >&2; exit 1; }
|
[ -f "$MODEL/config.json" ] || { echo "model not found: $MODEL" >&2; exit 1; }
|
||||||
# the mlx_lm of the harness venv if the repo is on this Mac, else the one on PATH
|
# the mlx_lm server: $MLX_SERVER, else ~/qwen-venv (mlx-lm 0.32.0, see docs/remote-model.md), else the harness venv, else PATH
|
||||||
if [ -x "$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" ]; then
|
for c in "$MLX_SERVER" "$HOME/qwen-venv/bin/mlx_lm.server" "$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" "$(command -v mlx_lm.server)"; do
|
||||||
SERVER="$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server"
|
if [ -n "$c" ] && [ -x "$c" ]; then SERVER="$c"; break; fi
|
||||||
elif command -v mlx_lm.server >/dev/null 2>&1; then
|
done
|
||||||
SERVER="$(command -v mlx_lm.server)"
|
[ -n "$SERVER" ] || { echo "mlx_lm.server not found: create ~/qwen-venv with mlx-lm 0.32.0 (docs/remote-model.md)" >&2; exit 1; }
|
||||||
else
|
# the baseline flags need mlx-lm 0.32.x; an older server does not know them
|
||||||
echo "mlx_lm.server not found (pip install mlx-lm==0.32.0)" >&2; exit 1
|
"$SERVER" --help 2>&1 | grep -q -- "--prompt-cache-bytes" || { echo "$SERVER is too old (no --prompt-cache-bytes). Install mlx-lm 0.32.0 in ~/qwen-venv (docs/remote-model.md)" >&2; exit 1; }
|
||||||
fi
|
echo "server: $SERVER"
|
||||||
"$SERVER" \
|
"$SERVER" \
|
||||||
--model "$MODEL" \
|
--model "$MODEL" \
|
||||||
--host 0.0.0.0 --port 8080 \
|
--host 0.0.0.0 --port 8080 \
|
||||||
|
|||||||
Reference in New Issue
Block a user