Developer Tools

Free AI Tool llama.cpp Just Got Faster on Your Own Graphics Card

⚡Run capable AI on your own PC — no subscription, no data leaving home.

Deep Dive

llama.cpp is free, open-source software that lets ordinary people run AI chatbots directly on their own laptop or desktop instead of paying a monthly fee to a cloud company. Its newest pre-release, version b11265, is a performance fix — not a flashy new feature. But it chips away at the biggest complaint about running AI at home: it's slower than the paid cloud versions. This update attacks that gap on one specific front.

The update concerns 'mixture of experts' models, which work like a large company where only the relevant specialists wake up for each question. Instead of activating the entire AI, just a few small slices do the thinking. That should make things fast and cheap. The problem was that llama.cpp was telling your graphics card to prepare for a huge crowd when only a handful of guests were coming — most of the card's worker threads sat idle. As the developer notes bluntly: 'Most workers in each group had nothing to do. This wasted time.' The sluggish step accounted for 55% of the whole job.

Why should you care? Because running AI on your own machine is the only way to get genuinely private AI. Your questions, documents and photos never leave your house. It's also free forever, with no per-message charges and no internet required. Every speed improvement like this one makes that option more realistic for regular people, not just hobbyists who enjoy fiddling with settings.

The honest catch: this is a developer pre-release, not a polished app. You still need some technical patience to install and use it, and most casual users won't notice this particular change yet. Still, the direction is clear. Big companies are racing to make AI run well on the hardware you already own — your laptop, your phone, your gaming PC. When that works smoothly, it changes who pays for AI, who controls it, and who gets to see your data.

Key Points
  • llama.cpp is a free tool that runs AI chatbots on your own computer, with nothing sent to a company's servers.
  • The fix targets 'mixture of experts' AI, which activates only a few specialist parts per question — the software was over-preparing, leaving most of your graphics card's workers idle.
  • One step that consumed 55% of total processing time is now much faster, meaning less waiting and lower compute cost for local AI.

Why It Matters

Faster local AI means private, subscription-free chatbots on hardware you already own — no cloud required.

📬 Get the top 10 AI stories daily