Developer Tools

llama.cpp b9785 released with hardened caps check and expanded platform support

The popular local LLM runner gets a security fix and broader GPU compatibility across devices.

Deep Dive

ggml-org's llama.cpp released version b9785 on June 25, with a hardened caps check (#24973). Builds are available for macOS (Apple Silicon, Apple Silicon with KleidiAI disabled, Intel), iOS, Linux (Ubuntu x64 and arm64 CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16, s390x), Android arm64, Windows (x64 and arm64 CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, HIP), and openEuler (x86 and aarch64 with 310p/910b ACL Graph). The repository has 118k stars and 19.9k forks. UI assets are included.

Key Points
  • llama.cpp b9785 includes a hardened caps check (#24973) to improve input safety in chat mode.
  • New builds for macOS Apple Silicon with KleidiAI, plus expanded Windows and Linux GPU backends (CUDA 12/13, ROCm 7.2, Vulkan, OpenVINO, SYCL, HIP).
  • Project now has 118k stars and 19.9k forks, reflecting massive community trust and adoption.

Why It Matters

Local LLM inference gets safer and more portable, enabling professionals to run models on any device without cloud dependency.

📬 Get the top 10 AI stories daily