AI Chatbots Still Hold Stereotypes—Even If They Don't Always Act on Them
AI models still judge you by your name or gender—but may not always let it change their advice.
Researchers tested whether AI chatbots let stereotypes about their users change what they actually say. First, they looked at smaller open-source models. The answer was clear: yes. When they nudged the models to think of a user as having a higher socioeconomic status, salary recommendations jumped by 141%. Men got higher salary suggestions than women, and women received less motivational language when asking whether they should apply for a job. Not a huge surprise—smaller models are known to absorb stereotypes from internet data.
The bigger question was about frontier models—the powerful, heavily fine-tuned chatbots from companies like OpenAI and Google. The researchers tested GPT-5.6, Gemini 3.1 Pro, and Claude Opus 5. All of them still formed strong stereotypes. Asked to invent fictional characters, they made surgeons, CEOs, and engineers mostly male, while nurses, teachers, and assistants were female. Race and class stereotypes showed up too: valedictorians were typically Asian, housekeepers were named "Maria," and a character named José was given an income of $42,000 while Wei earned $111,000.
But here's where it gets interesting. When those same fictional characters became real users asking for budgeting advice, all of them got the exact same answer—an assumed income of $4,000 per month. Gender made no difference either. A request signed "Emily" and one signed "Michael" received identical advice. The stereotypes existed, but they didn't always change the outcome. Only when a user explicitly said "I am a woman" did the advice become gendered.
What does that mean for you? The way you phrase a question to an AI matters as much as who you are. Stereotypes may be embedded, but cautious post-training in frontier models often keeps them from leaking into responses. Still, the stereotypes are easy to extract, and when they do surface—especially in smaller models or certain prompts—they can affect real decisions like salary advice, loan suggestions, or job guidance.
- Smaller AI models showed strong stereotypes: richer users got 141% higher salary advice, and women got less encouragement.
- Frontier models like GPT-5.6 and Claude Opus 5 still stereotype—typing 'José' earns $42k while 'Wei' earns $111k in stories.
- In practice, most frontier models gave identical advice no matter the user's name; only explicitly saying 'I am a woman' changed responses.
Why It Matters
AI bias can show up in real decisions like salaries or job advice, but outcomes depend on how you ask.