Developer Tools

Free App Lets You Run Powerful AI on Your Own Laptop

No subscription, no cloud, no one reading your chats — and it just got faster.

Deep Dive

A free, openly shared program called llama.cpp released a new version on September 17. The only real change in this update is faster support for a family of AI models made by Nvidia, called Nemotron. Specifically, it turns on something called "multi-token prediction" (MTP) — think of it like an AI finishing your sentence in its head before you finish typing, so words appear on screen faster. That's it. No flashy launch, no new product. Just a small speed upgrade to a tool used by millions of hobbyists and tinkerers.

So why should you care about a program you've never heard of? Because llama.cpp is the engine behind a quiet shift: running AI on your own machine instead of someone else's. When you use ChatGPT or Gemini, your words travel to a company's data center, and you pay a subscription or see ads. With llama.cpp, the whole AI lives on your laptop — it works on a plane with no Wi-Fi, costs nothing per question, and your private messages never leave your device. It's already used inside friendlier apps like Ollama and LM Studio, which wrap it in a simple point-and-click interface.

The honest catch: this is not plug-and-play. Setting up llama.cpp directly requires some technical comfort, and you need a reasonably powerful computer — a modern laptop, ideally with a decent graphics chip. This particular update also only helps one specific brand of AI model, so most people won't notice any difference at all. It's a small brick in a big wall, not a headline feature.

The bigger picture is where this is heading. Apple, Google, and Microsoft are all racing to put AI directly on phones and laptops, because it's cheaper for them, faster for you, and much better for privacy. Free open-source projects like this one are the reason that race is even possible. You probably won't install llama.cpp today. But the free, private, offline AI assistant you use in two years will likely owe it a thank-you.

Key Points
  • llama.cpp is free software that runs AI chatbots on your own laptop — no subscription and no internet required
  • This update speeds up Nvidia's Nemotron AI models using 'multi-token prediction' (the AI guesses several words ahead at once)
  • It's a small technical upgrade, but it's part of the bigger push to put AI on your device instead of in a company's data center

Why It Matters

Private, free AI on your own laptop keeps getting faster — your chats stay yours, and costs stay at zero.

📬 Get the top 10 AI stories daily