AI Knows the Difference Between 'Was Building' and 'Built' After All
You've heard AI is bad at nuance — turns out the test was broken, not the AI.
A new analysis challenges earlier claims that large language models fail at the imperfective paradox, arguing the benchmark itself was flawed. The researchers identify conceptual and evaluation mis-specifications, and show that models often do not affirm culmination but still accept simple-past hypotheses—a pattern they call "sufficiency bias." Prompting interventions shift label choices without reliably improving underlying semantic understanding. However, initial experiments with Qwen-7B, GPT-5.4, and Qwen-72B suggest aspectual classification can be context-sensitive, with performance comparable to human annotators. The takeaway: it may be a benchmark failure before a model failure.
- Previous AI grammar tests were flawed — 76% of key test questions were ambiguous, not clear failures.
- When tested with cleaner, matched sentence pairs, AI models like GPT-5.4 and Qwen matched human performance.
- This shows the importance of fair testing: bad benchmarks can mislead us about AI's real abilities.
Why It Matters
Fair AI testing shows language models understand nuance better than reported, so misjudging AI's abilities can mislead decisions in every field.