diff --git a/docs/remote-model.md b/docs/remote-model.md index e2538ca..9a2b4f8 100644 --- a/docs/remote-model.md +++ b/docs/remote-model.md @@ -25,24 +25,34 @@ stay on the **Mac mini**. The series is `python3 -m harness.localqwen` (see belo 31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors ``` (These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop. -3. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull` +3. **mlx-lm 0.32.0 on the MacBook** (the baseline flags need it; an older server answers `unrecognized arguments: --temp ...`). Check: + `~/qwen-venv/bin/python -c "import mlx_lm; print(mlx_lm.__version__)"` must print `0.32.0`. If there is no such venv: + ```sh + # with uv (recommended): brew install uv (once) + uv venv --python 3.11 ~/qwen-venv + uv pip install --python ~/qwen-venv/bin/python "mlx-lm==0.32.0" + # without uv (needs Python 3.10 or newer, check python3 --version; Homebrew: brew install python@3.11): + python3.11 -m venv ~/qwen-venv && ~/qwen-venv/bin/pip install "mlx-lm==0.32.0" + ``` + The start script uses `~/qwen-venv` first and refuses to start with an older server (clear message, nothing half started). +4. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull` (path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh` (192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`. The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB; only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w ` keeps the MacBook awake as long as the server lives. -4. **Start it** (a terminal window on the MacBook, leave it open; or in the background): +5. **Start it** (a terminal window on the MacBook, leave it open; or in the background): ```sh nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 & ``` (use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask "Do you want the application Python to accept incoming network connections?": **Allow**. -5. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and +6. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and `curl -s http://127.0.0.1:8080/v1/models` returns the model path. -6. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details). +7. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details). Both Macs must be in the same network (the mini is 192.168.178.x). -7. **Check from the Mac mini** (replace the IP): `curl -s http://:8080/v1/models`. It must show +8. **Check from the Mac mini** (replace the IP): `curl -s http://:8080/v1/models`. It must show `/Users/I301710/models/Qwen3.8-27B-4bit`. -8. **Start the series on the Mac mini:** +9. **Start the series on the Mac mini:** ```sh cd ~/projects/abap-llm/harness python3 -m harness.localqwen --base-url http://:8080/v1 --wait 600 diff --git a/train/serve_remote.sh b/train/serve_remote.sh index a7170a2..d1fed9b 100755 --- a/train/serve_remote.sh +++ b/train/serve_remote.sh @@ -4,14 +4,14 @@ # MacBook awake as long as the server process lives (caffeinate -w ). Usage: ~/serve_remote.sh Stop: kill $(cat ~/qwen_server.pid) MODEL="$HOME/models/Qwen3.8-27B-4bit" [ -f "$MODEL/config.json" ] || { echo "model not found: $MODEL" >&2; exit 1; } -# the mlx_lm of the harness venv if the repo is on this Mac, else the one on PATH -if [ -x "$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" ]; then - SERVER="$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" -elif command -v mlx_lm.server >/dev/null 2>&1; then - SERVER="$(command -v mlx_lm.server)" -else - echo "mlx_lm.server not found (pip install mlx-lm==0.32.0)" >&2; exit 1 -fi +# the mlx_lm server: $MLX_SERVER, else ~/qwen-venv (mlx-lm 0.32.0, see docs/remote-model.md), else the harness venv, else PATH +for c in "$MLX_SERVER" "$HOME/qwen-venv/bin/mlx_lm.server" "$HOME/projects/abap-llm/harness/train/.venv/bin/mlx_lm.server" "$(command -v mlx_lm.server)"; do + if [ -n "$c" ] && [ -x "$c" ]; then SERVER="$c"; break; fi +done +[ -n "$SERVER" ] || { echo "mlx_lm.server not found: create ~/qwen-venv with mlx-lm 0.32.0 (docs/remote-model.md)" >&2; exit 1; } +# the baseline flags need mlx-lm 0.32.x; an older server does not know them +"$SERVER" --help 2>&1 | grep -q -- "--prompt-cache-bytes" || { echo "$SERVER is too old (no --prompt-cache-bytes). Install mlx-lm 0.32.0 in ~/qwen-venv (docs/remote-model.md)" >&2; exit 1; } +echo "server: $SERVER" "$SERVER" \ --model "$MODEL" \ --host 0.0.0.0 --port 8080 \