Research & Papers

New Memory Trick Keeps AI Agents From Freezing — Wait Times Drop 75%

⚡Smarter memory cleanup means your AI assistant answers faster and costs less to run.

Deep Dive

When you talk to an AI agent (AI that can take actions for you, like booking a flight or fixing code), it re-reads the entire conversation every single turn. Servers save that reading work as "cached prefixes" — think of them as bookmarked notes. But when thousands of agents run at once, that memory fills up, and the server has to throw something away. This paper is about the rule it uses to decide what to delete.

The standard rule is LRU, short for "least recently used" — roughly, toss whatever hasn't been touched in the longest time. It sounds fair, but it isn't. Because agents all refresh their bookmarks at once, LRU ends up wiping out the entire history of just a handful of agents. When one of those unlucky agents comes back, it has to re-read everything from scratch, and the user sits there waiting.

RR-Evict does something different. It goes through every idle agent in turn — round-robin, like dealing cards — and deletes only the last small chunk of each one's notes. That way, many agents keep the useful beginning of their conversations, and whoever returns only re-reads a small missing piece. It doesn't need to predict which agent will come back next. Tested inside SGLang, an open-source engine that powers many AI products, the worst-case wait before the first word appears dropped by up to 75.4%, and repeated re-reading dropped by up to 65.7%.

The catch: this is a research paper, not a product you can buy. Big AI companies have to adopt it, and the gains vary by workload — someone chatting with a single AI won't notice a thing. Also, cheaper servers for AI companies don't automatically mean cheaper subscriptions for you.

Key Points
  • AI servers have limited short-term memory, and today's cleanup rule deletes a few users' entire chat histories while leaving others untouched.
  • The new rule, RR-Evict, deletes a small tail-end piece from every idle agent instead, so returning users wait far less.
  • In tests inside SGLang, the slowest response times fell up to 75% and wasted re-reading fell up to 66% — meaning faster, cheaper AI for everyone sharing the service.

Why It Matters

Faster, cheaper AI agents mean less waiting on chatbots and coding assistants, and potentially lower prices for tools you use daily.

📬 Get the top 10 AI stories daily