Research & Papers

New Research: AI Reading Scores Are Useful, But Don't Trust Them Blindly

The reading-level labels on books and health leaflets may be shakier than you think.

Deep Dive

If you've ever seen a book labeled 'Grade 5 reading level' or a health leaflet rewritten in 'plain English,' you've met readability assessment. It's the science of guessing how hard a piece of writing is to understand. In a new study posted to arXiv, researchers Joshua Wong and Chris Tanner compared two ways of doing that: old-fashioned formulas that count things like sentence length and word difficulty, and modern AI language models — the same kind of technology behind chatbots and translation tools.

They tested both across Arabic, English, French, Hindi and Russian. The AI models did recover the same signals the formulas rely on: how long sentences are, how complex the grammar is, and how varied the vocabulary is. That's useful news for anyone building tools that simplify text — writing assistants, translation apps, textbook publishers, and government agencies required to publish forms ordinary people can actually read.

But the study also carries a warning. The researchers point out that 'readability' labels are subjective — different human raters disagree about what counts as difficult. So an AI can score very well on a test while mostly learning surface patterns, like counting words, rather than truly grasping what makes writing confusing. They also found that models built for one language tracked the traditional formulas more closely than a single model trained on many languages at once.

The practical takeaway: readability scores are useful guides, not verdicts. If a tool tells you your email now reads at a sixth-grade level, that's a hint, not a guarantee that a real reader will find it easier. And when AI writes or rewrites text for you, a human check — ideally from an actual member of your audience — is still the best test of whether it makes sense.

Key Points
  • Old-fashioned readability formulas and AI models largely agree — but mostly on surface things like sentence length and vocabulary variety.
  • The study covered five languages: Arabic, English, French, Hindi and Russian, using a dataset called ReadMe++.
  • Because 'reading difficulty' is a human judgment, high AI accuracy doesn't prove the AI understands what actually confuses readers.

Why It Matters

Readability scores shape school books, health leaflets and AI writing tools — but they're a guide, not gospel.

📬 Get the top 10 AI stories daily