Files
abap-llm/docs/remote-model.md

88 lines
6.8 KiB
Markdown

# Local Qwen on the MacBook, harness on the Mac mini (series A)
Roles: the **MacBook only serves the model** (MLX, 4-bit). The harness, A4H, the MCP server, the records and the dashboard
stay on the **Mac mini**. The series is `python3 -m harness.localqwen` (see below). Written 2026-10-05.
## 1. Start the server on the MacBook (step by step)
1. **Plug the MacBook in and keep the lid open.** Close all heavy apps (the model needs about 16 GB, the prompt cache up to 6 GB).
2. **Check that this is the official baseline build** (same 4-bit weights as the baseline of 2026-10-04):
```sh
cd ~/models/Qwen3.8-27B-4bit
shasum -a 256 config.json model.safetensors.index.json tokenizer.json chat_template.jinja
```
The values must be:
```
14b65a0ee06517060a6bbd979bb1a8ff54e7b304b1a1f01d54344b88b8285e85 config.json
13b840162b4cb35c66fef7df072f7dbb4717908204364f5e5d9f9655a2758fa8 model.safetensors.index.json
06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523 tokenizer.json
c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041 chat_template.jinja
```
Full check (about 2 minutes, the weights): `shasum -a 256 model-0000*-of-00003.safetensors`
```
6cc1508e96fb5d0865dfd5753a79f4ec60651bf3e2a82844a7e8ae9c60528c0d model-00001-of-00003.safetensors
83f2a20ca8058f486a3634a27faf99587f4cd3c156a83dee34fb99e6ac178670 model-00002-of-00003.safetensors
31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors
```
(These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop.
3. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
(path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh`
(192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`.
The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB;
only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w <server pid>` keeps the MacBook awake as long as the server lives.
4. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
```sh
nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 &
```
(use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask
"Do you want the application Python to accept incoming network connections?": **Allow**.
5. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
`curl -s http://127.0.0.1:8080/v1/models` returns the model path.
6. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
Both Macs must be in the same network (the mini is 192.168.178.x).
7. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
`/Users/I301710/models/Qwen3.8-27B-4bit`.
8. **Start the series on the Mac mini:**
```sh
cd ~/projects/abap-llm/harness
python3 -m harness.localqwen --base-url http://<MacBook IP>:8080/v1 --wait 600
```
(`--wait 600`: waits up to 10 minutes for the server at the start. Run it with `nohup ... > runs/local_qwen.log 2>&1 &`.)
The 12 hour window starts with the first run, not with the command.
## 2. Stop
- **The series** (Mac mini): `pkill -f harness.localqwen` stops it, but the run in progress is cut and its objects stay in A4H. Better: wait for the
window to end (hard stop, clean) or ask Claude to stop it cleanly.
- **The server** (MacBook): `kill $(cat ~/qwen_server.pid)` (`caffeinate -w` ends when the server ends), or `pkill -f mlx_lm.server`.
Check: `curl -s -m 3 http://127.0.0.1:8080/v1/models` gives no answer.
- Closing the lid or a sleeping MacBook stops the server: the series then pauses by itself and resumes when the server answers (below).
## 3. What the harness does (series A)
- **Before the first run:** the server must answer, the served model path must end in `Qwen3.8-27B-4bit`, and a 16-token probe must work.
The answer is saved in `runs/local_qwen/preflight.json` (model path, speed of the probe). The weights themselves cannot be
checked remotely: the shasum comparison in step 2 is the proof of the same build.
- **Tasks and order** (`--plan-only` prints them): 3 x INTF, 3 x TABL, 3 x STRU, 3 x MSAG, 3 x exception, then 5 x DDLS (accepted
training tasks, no K variants, lowest ids). One task at a time.
- **Settings as the official baseline:** thinking off (`enable_thinking=false` per request), max_tokens 16384, loop guard 3,
tool budget 60 (CDS too), temperature 0.2, same system prompt, at most 80 turns.
- **Window:** 12 hours from the first run. At the end there is a hard stop, also in the middle of a run: the model request is dropped,
the run is not scored (no failure), its objects are deleted on A4H (the teardown needs about 20 s).
- **Server does not answer:** during a request the server is pinged every 30 s (3 misses in a row); between requests a failed
request is checked with a ping. Then the run ends cleanly (objects deleted, no scoring, folder `runs/local_qwen/_paused/`), the series
waits (a check every 30 s), and when the server answers again the same task starts again. The time of the pause counts in the 12 hours.
This is not counted as a failure.
- **Output** (`runs/local_qwen/`): `state.json` (the dashboard card), `results.jsonl` (one line per task), `summary.json` (at the end: pass rate and
failure types per kind), `accepted.jsonl` (accepted trajectories, **not for the first SFT**: `metadata.use_for_sft = false`, `series = local_qwen_A`),
`runs/` (one folder per run with `record.json`, trajectory, report). Resumable: start it again with the same arguments.
- **Acceptance** is the same filter as for the DeepSeek trajectories: score at least 80, end reason `report`, no harness error text.
Failure types per run: `loop`, `tool_budget`, `empty_response`, `model_error`, `not_active`, `contract`, `hidden_tests_not_run`,
`hidden_tests_failed`, `out_of_scope_changed`, `atc_priority1`, `release_syntax`, `low_score`.
- **Dashboard:** the card "Series A: local Qwen on the MacBook" in `runs/dashboard/index.html` (5-minute refresh).
## 4. Speed and cost
About 12 tokens per second with the 4-bit model (`train/README.md`); a task needs 20 to 40 minutes, so the 20 tasks take 7 to 13 hours:
at the slow end the window ends before the last DDLS tasks. No cloud cost: the ledger is not touched (the model is not a `:cloud` model).