Running a 27B model on two Arc Pro B60s
Two Intel Arc Pro B60s in a normal desktop now serve Qwen3.8-27B at about 53 tokens/s for a single user and around 250 tokens/s across eight, with a 262K context.
The setup
- 2× Arc Pro B60, 24GB each, in a consumer Alder Lake board. No PCIe switch.
- Qwen3.8-27B, split across both cards (tensor parallel,
-tp 2). - Intel's llm-scaler-vllm image, 0.26.0-b2 (vLLM 0.26.1).
Two cards were slower than one
First result: 13.7 tok/s on two cards, against about 24 on one. The kernel
won't allow GPU-to-GPU transfers on this chipset (it isn't on the
p2pdma allow-list), so every sync between the cards went through
system RAM. Worse, that path can't be captured in a graph, so the whole model
ran in slow eager mode.
The fix was piecewise graph capture with the all-reduce as a split point, plus MTP speculative decoding with the drafter kept out of the graph. That got it to 42 tok/s single and 172 at eight concurrent.
Turning on P2P anyway
The allow-list says "unknown", not "broken", so I wrote a small kernel module that allows this one host bridge. The catch: every GPU buffer then got mirrored into system RAM, 15–17GB per card, and the box ran out of memory in seconds. That turned out to be Intel's Level Zero runtime making allocations resident on the other card. Two environment variables fixed it.
| Through system RAM | P2P | |
|---|---|---|
| Time to first token, 206K prompt | 509 s | 151 s |
| Throughput, 8 users | 147 tok/s | 260 tok/s |
A watchdog switches back to the slower path if a card ever faults. It has done that twice in production, cleanly both times.
The slow leak
After a few hours, speed would drop from about 21 to 4 tok/s and the engine would eventually run out of memory. Intel's oneCCL library keeps a staging buffer for every message size it sees and never frees them, and real prompts come in a lot of sizes. Turning the cache off halved the speed. Rounding every message up to a fixed set of sizes fixed it: a two-hour soak of real agent traffic grew memory by 0.08GB.
FP8 or int4?
int4 is faster. FP8 solved more problems in a 24-problem LiveCodeBench run done as real agent tasks, mostly on the hard ones:
| 1 user | 8 users | LiveCodeBench (24) | |
|---|---|---|---|
| int4 | 86 tok/s | 342 tok/s | 12 |
| FP8 | 53 tok/s | 251 tok/s | 14 |
Every problem int4 solved, FP8 solved too. I went with FP8. A later search through community int4 builds found one that tied FP8 on a smaller test, but not convincingly enough to switch. I'm building my own int4 next.
Faster starts
The model ships with a vision encoder that vLLM loads and warms up on every
start. I send images to a separate, smaller vision model, so I turned it off
here (--language-model-only). Starts went from about 230 seconds to
131, and the KV cache grew 14%.
If you're picking a second card, check whether the board allows P2P first. Without it two cards still beat one, but long prompts and multi-user throughput suffer badly.