Claude Beats GPT in Smart New AI Reasoning Test
This could decide which AI you trust with complex questions.
A team of researchers created a new exam called SciReC to test how well AI chatbots can figure out relationships between different pieces of information. Think of it like a puzzle where you have to connect ideas, pictures, and sequences in the right way. This kind of "relational reasoning" is what we use every day to understand cause and effect, spot patterns, and make good choices.
The test is different from simpler AI quizzes because it involves back-and-forth conversation and images, not just text. The results show that Claude 4.6 did the best, answering 73% correctly, while GPT-5.4 got 68%. That might sound close, but the gap matters for people who rely on AI for research, planning, or learning. Even the best model still missed about a quarter of the questions, meaning there is plenty of room for error.
The researchers also looked at why AI gets things wrong. The biggest problem wasn't the images or the logic, but memory. AI often forgets earlier parts of a conversation, which is a huge issue for complex tasks like legal analysis, medical diagnosis, or tutoring. They also found that AI performs worst on astronomy questions and spatial tasks, but does best on psychology-related topics.
For everyday users, this means you shouldn't blindly trust AI when a task requires connecting many pieces of information over a long conversation. It's great for simple questions, but for serious decisions, you still need to double-check the details yourself. As AI gets better at memory and reasoning, these scores will improve, but this research shows exactly which skills still need work.
- Claude 4.6 outperformed GPT-5.4 on a new AI reasoning test (73% vs 68%).
- AI struggles most with remembering earlier parts of conversations, not just logic.
- The test found AI does worst on astronomy and spatial thinking, best on psychology.
- Even the top AI still gets one in four reasoning questions wrong.
Why It Matters
Knowing AI's weak spots helps you decide when to trust it and when to double-check.