88 lines
6.8 KiB
Markdown
88 lines
6.8 KiB
Markdown
# Local Qwen on the MacBook, harness on the Mac mini (series A)
|
|
|
|
Roles: the **MacBook only serves the model** (MLX, 4-bit). The harness, A4H, the MCP server, the records and the dashboard
|
|
stay on the **Mac mini**. The series is `python3 -m harness.localqwen` (see below). Written 2026-10-05.
|
|
|
|
## 1. Start the server on the MacBook (step by step)
|
|
|
|
1. **Plug the MacBook in and keep the lid open.** Close all heavy apps (the model needs about 16 GB, the prompt cache up to 6 GB).
|
|
2. **Check that this is the official baseline build** (same 4-bit weights as the baseline of 2026-10-04):
|
|
```sh
|
|
cd ~/models/Qwen3.8-27B-4bit
|
|
shasum -a 256 config.json model.safetensors.index.json tokenizer.json chat_template.jinja
|
|
```
|
|
The values must be:
|
|
```
|
|
14b65a0ee06517060a6bbd979bb1a8ff54e7b304b1a1f01d54344b88b8285e85 config.json
|
|
13b840162b4cb35c66fef7df072f7dbb4717908204364f5e5d9f9655a2758fa8 model.safetensors.index.json
|
|
06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523 tokenizer.json
|
|
c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041 chat_template.jinja
|
|
```
|
|
Full check (about 2 minutes, the weights): `shasum -a 256 model-0000*-of-00003.safetensors`
|
|
```
|
|
6cc1508e96fb5d0865dfd5753a79f4ec60651bf3e2a82844a7e8ae9c60528c0d model-00001-of-00003.safetensors
|
|
83f2a20ca8058f486a3634a27faf99587f4cd3c156a83dee34fb99e6ac178670 model-00002-of-00003.safetensors
|
|
31b8c91ef899f79efaaa69e3d2c096f6e2ebeb2ff20e29222abbd9ebc79e560a model-00003-of-00003.safetensors
|
|
```
|
|
(These are the hashes of `mlx-community/Qwen3.8-27B-4bit` on the Mac mini, `~/models/Qwen3.8-27B-4bit`.) If one differs, stop.
|
|
3. **Get the start script** `train/serve_remote.sh` onto the MacBook. Either, if the repo is on the MacBook, `cd ~/projects/abap-llm/harness && git pull`
|
|
(path `train/serve_remote.sh`), or copy it: `scp erhankeseli@192.168.178.40:~/projects/abap-llm/harness/train/serve_remote.sh ~/serve_remote.sh`
|
|
(192.168.178.40 is the Mac mini; use `.29` if that is its other address, check with `ifconfig` on the mini), then `chmod +x ~/serve_remote.sh`.
|
|
The script uses the same flags as the baseline (`train/serve.sh`): temp 0.2, top-p 0.95, top-k 20, min-p 0, max tokens 32768, prompt cache 4 / 6 GB;
|
|
only the host is `0.0.0.0` (port 8080), and `caffeinate -dimsu -w <server pid>` keeps the MacBook awake as long as the server lives.
|
|
4. **Start it** (a terminal window on the MacBook, leave it open; or in the background):
|
|
```sh
|
|
nohup ~/projects/abap-llm/harness/train/serve_remote.sh > ~/qwen_server.log 2>&1 &
|
|
```
|
|
(use `~/serve_remote.sh` if you copied it). The first start loads the model for about 1 minute. macOS may ask
|
|
"Do you want the application Python to accept incoming network connections?": **Allow**.
|
|
5. **Check on the MacBook:** `tail -f ~/qwen_server.log` shows `Starting httpd at 0.0.0.0 on port 8080`, and
|
|
`curl -s http://127.0.0.1:8080/v1/models` returns the model path.
|
|
6. **Find the MacBook IP** (the harness needs it): `ipconfig getifaddr en0` (Wi-Fi; if empty try `en1`, or look at System Settings > Wi-Fi > Details).
|
|
Both Macs must be in the same network (the mini is 192.168.178.x).
|
|
7. **Check from the Mac mini** (replace the IP): `curl -s http://<MacBook IP>:8080/v1/models`. It must show
|
|
`/Users/I301710/models/Qwen3.8-27B-4bit`.
|
|
8. **Start the series on the Mac mini:**
|
|
```sh
|
|
cd ~/projects/abap-llm/harness
|
|
python3 -m harness.localqwen --base-url http://<MacBook IP>:8080/v1 --wait 600
|
|
```
|
|
(`--wait 600`: waits up to 10 minutes for the server at the start. Run it with `nohup ... > runs/local_qwen.log 2>&1 &`.)
|
|
The 12 hour window starts with the first run, not with the command.
|
|
|
|
## 2. Stop
|
|
|
|
- **The series** (Mac mini): `pkill -f harness.localqwen` stops it, but the run in progress is cut and its objects stay in A4H. Better: wait for the
|
|
window to end (hard stop, clean) or ask Claude to stop it cleanly.
|
|
- **The server** (MacBook): `kill $(cat ~/qwen_server.pid)` (`caffeinate -w` ends when the server ends), or `pkill -f mlx_lm.server`.
|
|
Check: `curl -s -m 3 http://127.0.0.1:8080/v1/models` gives no answer.
|
|
- Closing the lid or a sleeping MacBook stops the server: the series then pauses by itself and resumes when the server answers (below).
|
|
|
|
## 3. What the harness does (series A)
|
|
|
|
- **Before the first run:** the server must answer, the served model path must end in `Qwen3.8-27B-4bit`, and a 16-token probe must work.
|
|
The answer is saved in `runs/local_qwen/preflight.json` (model path, speed of the probe). The weights themselves cannot be
|
|
checked remotely: the shasum comparison in step 2 is the proof of the same build.
|
|
- **Tasks and order** (`--plan-only` prints them): 3 x INTF, 3 x TABL, 3 x STRU, 3 x MSAG, 3 x exception, then 5 x DDLS (accepted
|
|
training tasks, no K variants, lowest ids). One task at a time.
|
|
- **Settings as the official baseline:** thinking off (`enable_thinking=false` per request), max_tokens 16384, loop guard 3,
|
|
tool budget 60 (CDS too), temperature 0.2, same system prompt, at most 80 turns.
|
|
- **Window:** 12 hours from the first run. At the end there is a hard stop, also in the middle of a run: the model request is dropped,
|
|
the run is not scored (no failure), its objects are deleted on A4H (the teardown needs about 20 s).
|
|
- **Server does not answer:** during a request the server is pinged every 30 s (3 misses in a row); between requests a failed
|
|
request is checked with a ping. Then the run ends cleanly (objects deleted, no scoring, folder `runs/local_qwen/_paused/`), the series
|
|
waits (a check every 30 s), and when the server answers again the same task starts again. The time of the pause counts in the 12 hours.
|
|
This is not counted as a failure.
|
|
- **Output** (`runs/local_qwen/`): `state.json` (the dashboard card), `results.jsonl` (one line per task), `summary.json` (at the end: pass rate and
|
|
failure types per kind), `accepted.jsonl` (accepted trajectories, **not for the first SFT**: `metadata.use_for_sft = false`, `series = local_qwen_A`),
|
|
`runs/` (one folder per run with `record.json`, trajectory, report). Resumable: start it again with the same arguments.
|
|
- **Acceptance** is the same filter as for the DeepSeek trajectories: score at least 80, end reason `report`, no harness error text.
|
|
Failure types per run: `loop`, `tool_budget`, `empty_response`, `model_error`, `not_active`, `contract`, `hidden_tests_not_run`,
|
|
`hidden_tests_failed`, `out_of_scope_changed`, `atc_priority1`, `release_syntax`, `low_score`.
|
|
- **Dashboard:** the card "Series A: local Qwen on the MacBook" in `runs/dashboard/index.html` (5-minute refresh).
|
|
|
|
## 4. Speed and cost
|
|
|
|
About 12 tokens per second with the 4-bit model (`train/README.md`); a task needs 20 to 40 minutes, so the 20 tasks take 7 to 13 hours:
|
|
at the slow end the window ends before the last DDLS tasks. No cloud cost: the ledger is not touched (the model is not a `:cloud` model).
|