Study: When AI Says 'Very Likely,' It Doesn't Mean What You Think
AI's confidence words don't match human ones — and that could mislead your decisions.
People describe confidence with words, not numbers. We say 'probably,' 'unlikely,' 'almost certain.' Psychologists have spent decades measuring what those words mean to humans — roughly how often 'likely' turns out to be true. This new paper, presented at a major machine learning conference, collected that human data and tested leading AI language models against it. The result: AI models attach noticeably different probabilities to the same words we use.
Why that matters is simple. You read AI output and unconsciously translate its words into risk. If a chatbot reviewing your symptoms says a condition is 'very likely,' you might picture 90% certainty. If the model's internal number is closer to 60%, you're making a decision on a false sense of confidence. The gap runs both directions — some phrases sound more cautious to us than the AI intends, others sound more certain.
The researchers also introduced a new way to measure this. Instead of asking a model the same question many times and averaging the answers — slow, costly, and noisy — they fit a mapping from each phrase to a probability that best explains when the model was actually correct. That lets you compare, word by word, what a phrase means to a human versus to an AI, and spot systematic patterns rather than one-off errors.
The catch: this is a research finding, not a fix. The mapping is specific to each model and topic, and it doesn't make an AI genuinely aware of its own limits. It also doesn't stop a model from sounding confident while being wrong. The practical takeaway for anyone using AI at work: treat phrases like 'very likely' as a rough vibe, not a calibrated risk score. Ask for numbers, sources, or a range — and verify anything that matters.
- AI models use words like 'likely' and 'possible' with different confidence levels than humans do, so the same phrase can mean two different things.
- The team's method maps each phrase to a real probability by checking how often the model was actually right — no repeated sampling needed.
- It's a warning, not a fix: don't treat AI hedging words as reliable risk estimates, especially for health, money, or legal questions.
Why It Matters
When AI says 'very likely,' don't assume it means what you'd mean — verify before acting.