Running a 27B model on two Arc Pro B60s
Two B60s in a normal desktop now serve Qwen3.8-27B at about 53 tokens/s for one user and 250 across eight, after a graph-capture change, a P2P kernel module and a oneCCL memory leak.
Notes from running local LLMs on Intel Arc cards: two Arc Pro B60s in a normal desktop, serving a 27B model to a few coding agents.
Two B60s in a normal desktop now serve Qwen3.8-27B at about 53 tokens/s for one user and 250 across eight, after a graph-capture change, a P2P kernel module and a oneCCL memory leak.