Study: Fancy AI Reading Tests Don't Make Better Recommendations
Researchers tested whether quizzing AI on 'reading evidence' helps it recommend — mostly no.
A team asked whether giving small AI models "reading comprehension" quizzes about evidence helps them choose a better recommendation interface. The answer, in their test, was no. They evaluated six small instruction-tuned checkpoints across four recommendation domains, with chronological evaluation and 3,426 evaluation users, ranking eight candidates per request. An interface chosen once on validation data for each domain and checkpoint scored 0.5524 NDCG@5, compared with 0.5447 for a baseline selector and 0.5428 for an augmented selector that added the diagnostic features. Adding those features changed NDCG@5 by -0.0019, with a 95% interval of [-0.0046, 0.0004] — an interval that includes zero and whose upper bound sits below the analysis plan's 0.005 improvement target. Matching the selectors' hyperparameters also left that upper bound below the target. Separately, evidence from retrieved similar users improved prompting by 0.0999 NDCG@5 over a control using randomly selected users matched for activity, and the evidence-reading tests revealed answer-position and tie-response biases. The authors note the results concern the tested selectors and candidate sets.
- Testing how well an AI 'reads evidence' did not improve its recommendations — the change was negative and statistically indistinguishable from zero.
- The one thing that helped a lot: feeding the model evidence from genuinely similar users, which beat random users by 0.0999 on a ranking-quality score.
- The tests also showed these models are biased by answer order and wording, so the same question can get different answers depending on how it's phrased.
Why It Matters
It's a reminder that trendy AI testing often adds cost without improving the suggestions you actually see.