Research & Papers

New Research: Hidden Settings Decide Whether AI Can Read Your Paperwork

⚡Wrong AI settings can mean wrong answers on rules that affect your money.

Deep Dive

Anyone building an AI assistant that answers questions about rules, contracts, or policy manuals faces the same chore: feed it documents, then let it look things up before answering. That lookup step is called RAG. This study took one such pipeline and tested it thoroughly. Three different document readers, three ways of chopping text into pieces, and five search models were combined and recombined, plus a keyword-search baseline, across four Indian government regulatory documents. That produced 72,000 scored results.

The headline finding is that there is no universal best setup. The document reader and the way you chop text affect each other significantly, so the winner changes from document to document. One widely used model, MPNet-base, performed poorly everywhere and failed badly on questions whose answers lived inside tables — a brutal weakness, since government and financial paperwork is full of tables. If you have ever had an AI chatbot confidently misread a fee schedule or a tax slab, this is likely why.

There is encouraging news buried in the numbers. The system preserved over 98% of the original evidence when converting documents, meaning almost no information was lost on the way in. So the failures are not about mangling your files; they are about ranking. The AI pulls up the right pile of text but puts the wrong page on top. That is a fixable search problem, not a broken-pipeline problem.

So what does this mean for you? Banks, insurers, law firms, hospitals, and government help desks are all racing to deploy these assistants. This paper hands them a ready-made 800-question test and open tooling, so they can check their setups instead of guessing. If you are buying or approving an AI document tool, ask two questions: how does it handle tables, and can we test it on our own documents first?

Key Points
  • A 54-way test on 800 questions from four Indian government rulebooks found no single best AI setup — results shift depending on how you split documents and which search model you use.
  • One popular search model, MPNet-base, failed noticeably, especially on questions answered inside tables, which are everywhere in official paperwork.
  • Nearly all source text (over 98%) survived the conversion step, so mistakes come from ranking the wrong passage first, not from losing information — a fixable problem.

Why It Matters

If your bank, insurer, or employer uses AI to answer policy questions, table handling and testing decide whether you get right answers.

📬 Get the top 10 AI stories daily