llama.cpp Update Lets Your Computer Run New AI Model
Run smarter AI on your own device — no cloud, no fees.
A new pre-release of llama.cpp is out: b10658, released on 27 Aug. The headline update is "spec: add DFlash2 support (local convolution + candidate selector)" with references to PRs #27342 and #27816. The release notes also include a long list of follow-up DFlash2 tweaks, like adding p_min, refactoring code, optimizing cost, and fixing CI. The release includes downloadable assets for macOS/iOS, Linux, Android, Windows, openEuler, and UI assets, and has already picked up 👍, 🎉, ❤️, and 🚀 reactions from users.
- llama.cpp lets you run AI models offline on your own hardware, keeping your data private.
- This update adds DFlash2, a new technique that could make AI faster and more capable on everyday devices.
- The code was partly written with help from another AI, Claude Opus 5, showing AI's growing role in development.
Why It Matters
You'll soon get faster, more private AI on devices you already own, without monthly fees.