Developer Tools

Ollama Update Makes Local AI Faster and Works Inside Claude

Your AI chats just got faster, smarter, and more private—right on your own computer.

Deep Dive

Ollama is like a personal AI engine for your computer. Instead of sending your questions to a company's server, it runs models like Llama or DeepSeek locally, which keeps your data private and works offline. This new version, v0.33.0, mainly improves the experience for people who use Claude Desktop. You can now pick which of your local Ollama models appear inside Claude, and even turn them on or off from the menu bar. Cloud models still show up, but only when you're signed in. It's like having a choice between your own assistant and a remote one, all in one place.

The big behind-the-scenes fix is about caching. Think of AI responding to a long conversation like reading a very long book. If you get interrupted and have to restart, you'd normally want to remember the page you were on. The new caching does exactly that. Before, if an AI request was cancelled, it might reprocess the entire conversation from scratch—imagine having to reread 46,000 of 47,000 tokens, which are basically chunks of text. Now it picks up right where it stopped, saving time and frustration. You'll notice faster responses after a hiccup, like a network drop or a pause.

The update also fixes a few annoying bugs. For example, a problem that caused the app to freeze when clients cancelled long prefill operations is gone. The developers also improved the setup screen for new users and made the app work better on Windows and Linux, not just Mac. If you use Claude, you'll no longer see a "tokens left" countdown that was actually slowing things down by messing with the cache.

Why should you care? If you use AI assistants for work or creativity, this means fewer slowdowns and more control over which models you use. And because Ollama runs on your machine, your private documents never leave your computer. That's a big deal for privacy-conscious people. It's not a flashy new feature you'll show off, but it's the kind of update that makes AI feel less like a fragile toy and more like a reliable tool.

Key Points
  • Turn your local Ollama AI models on or off directly inside Claude Desktop—no need to switch apps.
  • Smart caching means if an AI task gets interrupted, it resumes where it left off instead of starting from scratch—saving time on huge conversations.
  • Fixes a bug that could make AI reprocess 46,000 of 47,000 tokens, plus improves performance on Windows and Linux.

Why It Matters

Faster, more private, and more reliable AI on your own computer—less waiting, less cloud dependency, and fewer glitches.

📬 Get the top 10 AI stories daily