Research & Papers

Study: AI detectors flag honest edits but miss humanized AI text

64-80% false positives on light AI edits; under 4% detection after humanization.

Deep Dive

Researchers from Notre Dame and collaborators evaluated commercial AI detectors in a controlled study of published English abstracts across four domains, comparing pre-AI (2013-2015) and post-AI (2023-2025) writing. They found that light, guideline-compliant AI editing—just refining an abstract—is flagged as AI-generated 64-80% of the time by tools like Pangram and GPTZero. Meanwhile, unmodified recent abstracts—written entirely by humans alongside AI prevalence—are flagged at 9-15%, with non-STEM fields flagged at rates significantly higher than STEM (p<0.001).

The false positive pattern tracks long-token density and Academic Word List usage, not actual authorship intent. Worse, when the same AI-edited texts were run through Undetectable AI's humanization tool, evasion was near-total: fewer than 4% of AI-labeled rewrites remained flagged (a >96% false negative rate). This creates a perverse incentive: honest students who use AI for light editing face higher sanction risk than those who deliberately obfuscate AI-generated content. The authors conclude detector scores should not serve as standalone misconduct evidence and call for policy reform.

Key Points
  • Light AI editing of abstracts flagged as misconduct 64-80% of the time by GPTZero/Pangram
  • Unmodified 2023-2025 human abstracts flagged 9-15%, with non-STEM rates far higher than STEM
  • Undetectable AI humanization reduces detection to under 4%, making evasion easier than honest editing

Why It Matters

AI detectors punish honest AI-assistance while enabling cheaters—universities need evidence-based policies, not score-based accusations.

📬 Get the top 10 AI stories daily