Study: AI detectors flag honest edits but miss humanized AI text
64-80% false positives on light AI edits; under 4% detection after humanization.
Researchers from Notre Dame and collaborators evaluated commercial AI detectors in a controlled study of published English abstracts across four domains, comparing pre-AI (2013-2015) and post-AI (2023-2025) writing. They found that light, guideline-compliant AI editing—just refining an abstract—is flagged as AI-generated 64-80% of the time by tools like Pangram and GPTZero. Meanwhile, unmodified recent abstracts—written entirely by humans alongside AI prevalence—are flagged at 9-15%, with non-STEM fields flagged at rates significantly higher than STEM (p<0.001).
The false positive pattern tracks long-token density and Academic Word List usage, not actual authorship intent. Worse, when the same AI-edited texts were run through Undetectable AI's humanization tool, evasion was near-total: fewer than 4% of AI-labeled rewrites remained flagged (a >96% false negative rate). This creates a perverse incentive: honest students who use AI for light editing face higher sanction risk than those who deliberately obfuscate AI-generated content. The authors conclude detector scores should not serve as standalone misconduct evidence and call for policy reform.
- Light AI editing of abstracts flagged as misconduct 64-80% of the time by GPTZero/Pangram
- Unmodified 2023-2025 human abstracts flagged 9-15%, with non-STEM rates far higher than STEM
- Undetectable AI humanization reduces detection to under 4%, making evasion easier than honest editing
Why It Matters
AI detectors punish honest AI-assistance while enabling cheaters—universities need evidence-based policies, not score-based accusations.