Research & Papers

Stanford's Purlin Makes AI Chips Talk 5x Faster

⚡Faster chatbots and cheaper AI — by fixing how chips gossip with each other.

Deep Dive

When you type a question into an AI chatbot, one computer doesn't answer it. The work is spread across thousands of specialized chips that pass partial answers back and forth millions of times per second. That constant chatter is called "collective communication," and it's often the slowest part of the whole operation. A Stanford team — Osayamen Jonathan Aimuyo, Swapnil Gandhi, and Christos Kozyrakis — built a framework called Purlin to fix it.

The problem, they argue, is that today's systems tangle three things together: what the data means, who decides when and where it moves (the "orchestration"), and how it physically moves (the "datapath"). It's like a kitchen where the recipe, the head chef's orders, and the cooks' hands are all fused into one machine — change anything and everything breaks. Purlin separates them. You describe the data layout and the operation, a coordination protocol called SNAC figures out the timing, and a low-level layer called Atom does the actual copying and combining on the chip. That means new hardware can be plugged in without rewriting everything above it.

The results are real. Tested on Nvidia's A100, H200, and B200 chips, Purlin made seven common communication patterns up to 5.14 times faster and moved up to 4.5 times more data per second. Dropped into SGLang, a widely used tool for running AI models, it improved offline serving throughput by about 13% on average and boosted live chatbot responsiveness by 26% on average — up to 2.85x when the system was overloaded, exactly when users get the most frustrated. Image generation with diffusion models also got about 13% faster end to end.

So what's the catch? This is a research paper, not a product you can download and use today. It targets cutting-edge data-center GPUs and needs software changes to adopt. The headline percentages may sound modest in everyday terms, but in data centers they translate directly into serving more users on the same expensive hardware — which is how AI gets cheaper and snappier for everyone.

Key Points
  • Purlin splits the 'who decides' part of AI chip communication from the 'how it moves' part, so new hardware can be swapped in without rebuilding everything.
  • Tested on Nvidia A100, H200, and B200 chips, it made chip-to-chip data exchange up to 5.14x faster and 4.5x more bandwidth-efficient.
  • Inside the popular SGLang chatbot tool, it made live AI responses up to 2.85x more responsive during peak overload — exactly when things usually slow to a crawl.

Why It Matters

Faster chip chatter means snappier chatbots, quicker image generation, and cheaper AI — often without buying new hardware.

📬 Get the top 10 AI stories daily