Researchers Found AI Chatbots Sharing Notes — And Mixing Them Up
A cost-saving shortcut inside AI systems can quietly give you wrong answers and leak your prompts.
Every time you ask an AI chatbot something, the system writes itself a kind of scratch note — the words you typed, already processed — so it doesn't have to redo that work for the next part of the answer. This is called a "KV cache" (the AI's short-term memory of your conversation). Because that work is expensive, AI companies now pool these notes across many users and many different models, the way a busy kitchen pre-chops vegetables for every dish on the menu. It makes answers faster and cheaper.
The problem, according to a new paper from researchers Wei Song, Yuxin Cao, Xi Zheng, Leo Zhang and Xiao Cheng, is that the shared pool labels each note by the words in it but not by which model, settings, or customer it came from. So notes written by one AI can be handed to a completely different one. In tests covering 12 models from seven families and more than 160 configurations, using the official vLLM and SGLang software that many companies rely on, this mix-up dropped answer accuracy from 94% to 64%. When the mismatched memory was especially incompatible, reasoning accuracy fell to zero.
There's a privacy angle too. AI systems normally add random noise — called "salt" — to hide how long they take to respond. When that noise is missing, an outsider who simply measures response timing could identify 93% of what a user had typed. That means your prompt, not just the answer, can leak.
The fix is an idea the authors call a "provenance contract": a permanent label that travels with each note, recording where it came from and under what rules it was made. Any mismatch, and the note is thrown out instead of reused. Implemented in vLLM and SGLang, it removed the bad reuse while keeping the legitimate sharing, and added less than 0.34 milliseconds of delay — far less than normal run-to-run variation. This isn't a reported real-world breach; it's a structural flaw the authors say was hiding in plain sight. If your company builds or buys AI, it's worth asking which version you're running.
- "KV cache" is an AI's scratch notes for your conversation — companies pool them across models to save money and speed up replies
- Mixing notes between different AI models cut accuracy from 94% to 64%, and one test fell to zero
- Without random timing noise, 93% of a user's prompt could be identified just by watching response speed
Why It Matters
A cheap shared-memory shortcut inside AI could silently give your business wrong answers or expose what customers typed.