AI Safety

LLMs like GPT-4 and Claude show negative bias toward intellectual disabilities

Five major AI models consistently depict people with ID as younger, dependent, and needing help.

Deep Dive

Researchers from Oregon State University and collaborators used five LLMs—GPT-4-Turbo, GPT-4o, Meta Llama-3-3-70B-Instruct, Anthropic Claude-3-5-Sonnet, and Mistral-Large-2411—to generate over 25,000 stories from 10 prompt stems. Each stem was tested with and without descriptors for intellectual disabilities (ID). A separate GPT-4-Turbo instance analyzed the stories for representational differences tied to known bias themes from literature.

The findings showed consistent negative implicit biases beyond clinical characteristics of ID. Models portrayed people with ID as younger (infantilization), more inspirational/symbolic, dependent, needing help or rescue, and prompted hesitation or negative perception. These patterns echo historical discrimination and underline the risk that LLMs may reinforce harmful stereotypes in decision-making contexts like hiring, healthcare, or social services.

Key Points
  • Study analyzed 25,000 AI-generated stories from 5 LLMs (GPT-4-Turbo, GPT-4o, Llama-3-70B, Claude-3-5-Sonnet, Mistral-Large-2411) comparing prompts with and without ID descriptors.
  • Detected biases included infantilization, paternalism, depictions as inspirational/symbolic, dependency, and negative perception or hesitation.
  • Bias patterns were consistent across models and reflect historical discrimination against people with intellectual disabilities.

Why It Matters

AI models risk perpetuating harmful stereotypes about people with intellectual disabilities, demanding proactive bias auditing and mitigation.

📬 Get the top 10 AI stories daily