Llama AI Now Runs Faster on Your Phone's Graphics Chip
Your phone could soon run powerful AI without an internet connection.
The llama.cpp project added f16 support to the fill/set_rows operations in its webgpu backend, per pull request #29897. That's the entire change described. The rest of the page is a build-attestation list — macOS/iOS (including macOS Apple Silicon, macOS Intel disabled, iOS XCFramework), Linux builds across CPU, Vulkan, CUDA 12, CUDA 13, ROCm 10.0, OpenVINO, SYCL FP32, SYCL FP16, and Linux arm64 Snapdragon (CPU, Adreno GPU, Hexagon NPU), Android arm64 (CPU and Snapdragon CPU/Adreno GPU/Hexagon NPU), Windows x64 and arm64 builds across CPU, OpenCL Adreno, CUDA 12, CUDA 13, Vulkan, OpenVINO, SYCL, and ROCm 10.0, openEuler x86 and aarch64 (310p, and 910b with ACL Graph), plus a UI build.
- Llama.cpp added f16 support to WebGPU, making AI run faster on your device's graphics chip.
- This means AI apps could be quicker and use less battery when running locally on your phone or computer.
- You'll need a newer device that supports f16 in WebGPU; it's still experimental.
Why It Matters
Faster, more private AI on your phone without cloud dependence.