Intel Arc B140 build runs Llama.cpp on Xeon W-2255
A $10K Intel Arc B140 workstation runs Llama.cpp at 64GB VRAM...
Deep Dive
A local inference build from /u/mazarax: ASUS WS C422 PRO/SE motherboard, 10-core Xeon W-2255, 64GB ECC RAM, and 64GB VRAM. The pimped case features TurboLEDz that indicate the frequencies of the 10 Xeon cores. It runs llama.cpp with a SYCL backend, with Khronos-stack and MESA stack built from git sources on Ubuntu 26.04.
Key Points
- Intel Arc B140 GPU drives a $10K workstation with 64 GB VRAM and Xeon W-2255 CPU
- SYCL-backed Llama.cpp runs on Ubuntu 26.04 with bleeding-edge Khronos & MESA stacks from Git
- TurboLED core-frequency display adds DIY flair while benchmarking LLM inference locally
Why It Matters
Proves Intelβs Arc silicon can rival GPUs for local LLM workloads, cutting cloud costs for power users.