Kog claims 30x faster LLM inference on standard GPUs
French startup hits 3,000 tokens/sec on Nvidia and AMD chips
French startup Kog is betting that conventional GPUs still hold untapped inference power. In May, its tech preview went viral on Hacker News, demonstrating “extremely fast single-request decoding” on standard enterprise hardware like AMD MI300X and Nvidia H200. The demo achieved an impressive 3,000 tokens per second, but only with a purpose-built 2B-parameter model called Laneformer 2B. CEO Gaël Delalleau says the approach can scale to large LLMs, promising a 30x speedup over current inference stacks — a claim skeptics question given the massive memory bandwidth demands of bigger models.
Based on early feedback, Delalleau expects software engineering to be the first commercial use case, targeting professionals who pay premium prices for Claude's Fast Mode due to multi-hour waits. Kog also has design partners building prompt-to-game and prompt-to-app tools, where faster inference directly boosts revenue. The startup has pivoted to focus on larger models after learning customers won't fine-tune small ones. Delalleau, a solid-state physicist and white-hat hacker, brings a reverse-engineering mindset to GPU engineering, treating each new chip as a puzzle to solve — though this hands-on approach limits Kog to a few GPUs per year with its 11-person team. The startup is backed by Scaleway, Bpifrance, and French Tech 2030, positioning itself within Europe's push for AI sovereignty.
- Kog's inference engine hit 3,000 tokens/sec on a 2B-parameter model using Nvidia H200 and AMD MI300X GPUs
- Company claims 30x faster LLM inference is achievable, but has yet to demo it on large-scale models
- CEO Gaël Delalleau's background in physics and cyber security drives a low-level reverse-engineering approach to GPU optimization
Why It Matters
Faster software-only inference on existing GPUs could slash AI costs and latency for enterprises without new hardware investments.