Research & Papers

New AI Report Card Grades Search Results More Fairly

Google, Bing, and AI chatbots may soon show you better answers — thanks to a fairer report card for search results.

Deep Dive

A new evaluation metric called Rank-Deviation Quality (RDQ) is introduced for retrieval and ranking systems, designed to handle queries with anything from a single correct answer to many valid results. RDQ scores a candidate ranking against an ordered reference list, applying a rank-deviation penalty to each retrieved item based on its output position, and giving zero credit to items outside the reference list. Unlike metrics that need absolute relevance grades, RDQ uses ordinal rankings, and unlike rank-correlation measures like Kendall's tau, it accounts for both which items are returned and how they are ordered. In testing, RDQ achieved the highest median empirical power at 100 on a 5,000-query point-of-interest dataset among 13 metric configurations. On TREC Deep Learning benchmarks, RDQ reached comparable median power at 25 queries, while NDCG was higher at 100.

Key Points
  • Researchers created RDQ, a new way to grade AI search results that cares about both accuracy and order
  • Early tests on 5,000 real queries show it performs better than older grading methods
  • This could mean search engines and AI assistants finally show you the best answer first, not just any answer

Why It Matters

Better search rankings could save you minutes every day by putting the right answer at the top.

📬 Get the top 10 AI stories daily