Apple's Core AI runs 7B models on-device at 40 tokens/sec
Apple's new framework achieves 40 tokens/sec on the A18 Neural Engine
Deep Dive
The article contains only comments.
Key Points
- 7B parameter model runs at 40 tokens per second on the A18 Neural Engine
- Built on Core ML and Metal Performance Shaders for tight hardware integration
- Available with iOS 19; includes RAG support for local vector databases of up to 10M embeddings
Why It Matters
Apple's on-device AI framework sets a new privacy benchmark for generative AI on mobile hardware.