New Research Makes AI Assistants Up to 3.7x Faster Using Cheap Storage
Faster AI replies could mean less waiting and cheaper AI subscriptions for everyone.
When you chat with an AI assistant that can actually do things — book a table, read your files, search the web — it has to remember everything that happened in earlier steps. That memory is called a KV cache, and keeping it around is what lets the AI avoid re-reading the whole conversation each time. The problem: that memory is expensive. Companies normally park it in fast, pricey computer memory (DRAM), and as conversations stretch on, the AI gets slower and the bills get bigger.
A research team from China has a different idea, published as a paper called Janus. Instead of expensive memory, they store the AI's conversational memory on SSDs — the same kind of fast storage drive in a laptop, much cheaper than RAM. The catch with SSDs is they're slower to read from, so naively putting them in the loop would make things worse. Janus solves that by having the AI guess what it will need next, based on earlier clues, and start loading it before it's actually needed. It's like a restaurant that starts cooking your order before you sit down, because the host already guessed what you'd want.
The team also cleaned up how the data is stored — grouping scattered pieces together and pausing background chores while reading — so the drive isn't tripping over itself. Across three AI models and three realistic assistant scenarios, Janus started responding up to 3.7 times faster, averaging roughly 1.2 to 1.85 times faster, without hurting the quality of the answers.
Why should you care? Two reasons. First, latency: if you use AI agents for work, every second of waiting adds up across dozens of tasks a day. Second, cost: cheaper memory means AI companies can serve more users on the same hardware, which tends to push subscription prices down over time. The honest caveat is that this is a research paper, not a shipped product — it hasn't been tested at the scale of ChatGPT or Gemini, and real-world gains often shrink once you leave the lab.
- Janus stores an AI's memory of past conversation on cheap SSD drives instead of expensive RAM — like keeping your notes in a filing cabinet rather than on your desk.
- It predicts what the AI will need next and starts fetching it early, cutting waiting time by up to 3.7x in tests.
- Cheaper memory could eventually mean lower AI subscription costs and snappier tool-using assistants — but this is lab research, not yet a product.
Why It Matters
Faster, cheaper AI assistants mean less waiting on tasks and potentially lower prices for the tools you already use.