Why AI Gets Financial Numbers Wrong — And the Fix That Works
A small ranking tweak could cut wrong answers when AI reads your financial documents.
When you ask an AI chatbot about a company's finances, it doesn't memorize every report. Instead it looks things up first — a technique called RAG (letting AI search documents before answering). It scans thousands of passages and sends only a short list, maybe the top ten, to the AI that writes the answer. That short list is the whole ballgame. If the right passage isn't on it, the AI simply can't get the number right.
Here's the problem the researchers identified. Financial documents are full of passages that sound relevant but are wrong in a crucial way: they mention the topic you asked about, but from a different year, a different business segment, or a different accounting metric. The ranking system can't tell the difference, so these decoys fill up the limited short list and shove the genuinely correct evidence out. The researchers call this "financial evidence crowding." In their tests, mismatched passages cut the share of correct evidence reaching the AI's short list by nearly 15 percentage points.
Their fix, called FinDeCrowd-RAG, is essentially a smarter tiebreaker. It adds a second layer of scoring that checks whether a passage actually fits the question's specific constraints — the right period, segment, or metric — instead of just matching keywords. A built-in gate decides when to apply this adjustment and when to leave the original ranking alone, so it can't make things worse. In controlled tests, correct evidence in the top ten jumped from 76% to 90%. On a more realistic index, it rose from 42% to 49%. With the same AI writer, answer accuracy and citation quality improved on FinanceBench and FinQA, two widely used financial question sets.
The catch: the fix only rescues evidence that the first search already found — it can't invent what wasn't there. And it was tested on financial questions, not on your email or medical records. Still, as more people use AI to read annual reports, earnings calls, and investment research, getting the right passage into the AI's hands may matter more than making the AI itself smarter.
- AI gets financial questions wrong not because it's dumb, but because its search returns look-alike passages about the wrong year or metric.
- The researchers' ranking fix lifted correct evidence in the AI's short list from 76% to 90% in controlled tests, and from 42% to 49% in more realistic ones.
- Better evidence made answers more accurate on FinanceBench and FinQA — two standard financial question sets used to grade AI.
Why It Matters
Fewer wrong numbers when AI reads financial reports, benefiting investors, analysts, and anyone checking a company's books.