Developer Tools

Free AI Software llama.cpp Now Runs 3x Faster on Intel Graphics

⚡If you run AI on your own PC, your waiting time just got much shorter.

Deep Dive

A volunteer-run, free project called llama.cpp lets ordinary people run AI chatbots directly on their own laptop or desktop — no subscription, no internet connection, and nothing sent to a company's servers. This week it published a small update, mostly plumbing: it switched to Intel's newest software toolkit, called oneAPI 2026.1. That sounds dull, but the payoff is real. Intel's older toolkit was quietly dropping support for oneDNN, a piece of math software that speeds up AI work, so the project moved early rather than lose the speed later.

The numbers are the interesting part. On an Intel Arc B570 graphics card, the AI's "reading" step — the part where it chews through your question and any documents you attached before it starts answering — jumped from 434 to 1,331 units per second. That is about 3.1 times faster. The "writing" step, where it types out the reply, improved only slightly, from roughly 46 to 50 units per second. In practice: the pause before the first word appears shrinks dramatically, especially on long documents.

Why should you care? Local AI is the privacy-friendly, subscription-free option. Your emails, medical notes and contracts never leave your machine. The biggest complaint people have about it is that it feels slow. Cutting the reading time by two-thirds turns a coffee-break wait into a few seconds, which makes the free route genuinely usable for everyday tasks like summarizing a report or searching your own notes.

The catch: this only helps people with newer Intel graphics chips. If your computer runs Nvidia, AMD or Apple silicon, nothing changes for you this week. It is also a pre-release build, and installing local AI still takes some patience — it is not yet a double-click app. And since Intel plans to drop that oneDNN speedup in 2027, today's gain buys time, not permanence.

Key Points
  • llama.cpp is free software that runs AI chatbots on your own computer, so nothing you type is sent to a company.
  • On an Intel Arc B570 graphics card, the AI reads your question about 3.1 times faster — from 434 to 1,331 units per second.
  • The speed-up only applies to newer Intel graphics chips; Nvidia, AMD and Apple users see no change, and it is still a test build.

Why It Matters

Free, private AI on your own laptop just became fast enough to actually use for long documents.

📬 Get the top 10 AI stories daily