Want Better AI Answers? Make It Double-Check Itself, Study Finds
Smarter searching barely helped, but self-checking made answers noticeably better.
When AI answers questions about your documents — a setup called RAG (letting AI look things up before answering) — engineers usually pile on clever shortcuts. Some expand the section around a matching paragraph. Some send an "agent" (AI that takes actions on its own) to search repeatedly. Some read the table of contents first. This paper tested all 16 on/off combinations of four such features, across 768 separate test conditions, on five public documents ranging from 78 to 492 pages. Every answer was graded against a verified correct reference.
The surprise: the feature that mattered most wasn't about finding information at all. It was a completeness check — after the AI writes its answer, it scans for gaps and tries to fill them. On its own, this scored 4.31 out of 5, beating every combination that left it out, including the three-feature combo of section expansion, agentic search, and table-of-contents retrieval at 4.11. Reading the table of contents first also helped, and cost nothing extra in AI usage. Fancy agentic search was a wash — helping some questions and hurting others.
The catch is speed. The very best configuration roughly doubled how long you wait for an answer, creating a real trade-off between quality and patience. The study also found that which feature looks best depends heavily on the kind of question you ask. Test with only one question type, and you'll rank the features wrong — the completeness check's advantage jumped to a huge effect on questions that demanded thorough coverage. So if you hear a company claim its search feature is best, ask what kinds of questions they tested.
The practical lesson for anyone buying or building AI tools: spend your effort on making the AI verify its own work, not on ever-fancier searching. It's like a research assistant who reads everything but forgets to proofread — versus one who checks her draft against the assignment. The second one is who you actually want.
- Having AI review its own answer after writing it beat every other improvement tested — including fancier search methods.
- A simple table-of-contents trick gave a solid accuracy boost at zero extra AI cost; AI-driven repeated searching was inconsistent.
- The best setup roughly doubled wait times, so expect a real trade-off between answer quality and speed.
Why It Matters
If you use AI to research documents, demand answer-checking, not just better search — and expect to wait longer for it.