DeepSeek's New AI Cuts Costs by Shrinking Memory Use
Cheaper AI means cheaper products for you — and less strain on power-hungry data centers.
DeepSeek, the Chinese AI company that shook up the industry with cheap, capable models, launched V4.1-Flash on Thursday. The trick is how it's built. Rather than firing up all 552 billion of its internal parts for every question, it only wakes up the handful it needs — about 8 billion — like a giant firm where only a few specialists handle your case instead of the whole company. It can also hold roughly a million words of context at once, enough to keep an entire book or months of chat history in mind.
The bigger news is memory. Think of an AI as a waiter taking your order: it has to remember everything said so far, and that notepad costs real money to keep around. DeepSeek says the notepad just got much smaller — from about 3,514 bytes per word of conversation down to 890, a quarter of the memory chips and an eighth of the storage. For businesses running millions of AI requests, that means serving more people at once without buying more machines. Off-peak prices can now fall as low as 0.02 yuan per million words processed, a fraction of what top models cost.
There's a catch worth naming. All the impressive test scores — including beating OpenAI's GPT-5.6 on a coding test and Anthropic's Claude Opus 5 on another — come from DeepSeek's own testing. Independent labs haven't reproduced them, and a score on a test isn't the same as being useful in your actual job. Claude still wins on other measures. On the plus side, DeepSeek posted the model's underlying code publicly for anyone to check and use freely.
Finally, DeepSeek is doing something unusual to its own products. Starting September 14, requests to its older, more expensive V4 Pro model get quietly rerouted to V4.1-Flash and billed at the cheaper rate. A company executive said on X that it "would not be appropriate" to keep giving users a weaker model at a higher price. If that pricing holds, expect rivals to cut prices too.
- DeepSeek's new V4.1-Flash uses a quarter of the expensive memory chips and an eighth of the storage of the previous version, so running AI gets cheaper.
- Off-peak use can cost as little as 0.02 yuan per million words processed — a big drop that could pull down prices across the industry.
- DeepSeek's own tests show it beating OpenAI's GPT-5.6 on coding tasks, but those numbers haven't been checked by independent outsiders.
Why It Matters
Cheaper, lighter AI means lower prices for AI tools you use, and less strain on power and hardware.