Different AI Models Give Different Answers — So Can You Trust AI Assessments?
If AI can't give the same answer twice, can we trust it to grade or interview you?
AI isn't just a tool for answering trivia questions. Schools, companies, and governments are exploring how to use it to evaluate people — from grading essays to interviewing job candidates. A new study shows that might be trickier than it sounds. Researcher Jiangang Hao took real conversations from collaborative chat sessions and used different AI models to generate replies. He compared the answers to see if they were similar.
The results were clear: the AI model you choose matters a lot. Swap one model for another and the replies change, even when given the same conversation. Adding previous chat history also changed the answers. That means the same 'conversation' can lead to very different responses depending on which AI you happen to use.
Why should you care? Because assessment tools make decisions that affect people's lives. If two students get the same AI-graded interview, but the AI models were different versions, their scores could differ for no reason other than the software. The researcher argues that prompting tricks alone can't fix this. We need better 'infrastructure and design strategies' to make AI responses more consistent and fair.
The practical takeaway for ordinary people? Don't assume AI is objective. If you're being evaluated by an AI, ask about consistency and oversight. If you're a manager, be cautious about automating hiring or performance reviews. As AI models keep evolving, we need rules and testing to keep them fair — just like we have for human-made tests.
- Different AI models produce different replies in the same conversation, and chat history changes the answers too.
- This inconsistency can make AI-powered grading, interviews, and assessments unfair or unreliable.
- The researcher recommends better testing and design standards to keep AI responses stable as models constantly evolve.
Why It Matters
AI-powered grading or interviews may yield different results depending on the model, risking unfair decisions for students and job seekers.