Developer Tools

Apple's Core AI runs 7B models on-device at 40 tokens/sec

Apple's new framework achieves 40 tokens/sec on the A18 Neural Engine

Deep Dive

The article contains only comments.

Key Points
  • 7B parameter model runs at 40 tokens per second on the A18 Neural Engine
  • Built on Core ML and Metal Performance Shaders for tight hardware integration
  • Available with iOS 19; includes RAG support for local vector databases of up to 10M embeddings

Why It Matters

Apple's on-device AI framework sets a new privacy benchmark for generative AI on mobile hardware.

📬 Get the top 10 AI stories daily