Research & Papers

New AI Test Shows Chatbots Still Fail at Questions Mixing Text and Images

Could expose which AI assistants actually understand what they're looking at.

Deep Dive

AI chatbots like ChatGPT are getting smarter at answering questions. But the smartest ones still make big mistakes when the answer requires understanding both words and pictures. This new study from university researchers introduces a test called CrossModalQA that measures exactly that ability. Think of it like a quiz where every question sends you on a mini treasure hunt across written articles and photos, and you have to connect clues from both to reach the answer.

The test is built from nearly 5,000 Wikipedia pages and more than 4,400 images from Wikimedia Commons. It contains 1,863 questions that are not simple. Many ask you to read an article, find a photo, then read another article, then compare several images. For example, to answer a question about a historic event, you might need to identify a landmark in one photo, look up the architect in text, and then find a matching image of a similar bridge. That is called multi-hop reasoning, because it takes many small leaps.

The researchers found something striking: the best current AI systems often struggle with these questions. When they tried to search for evidence across text and images, they frequently recovered only part of the story. Worse, adding incomplete search results could mislead the AI and hurt its accuracy more than if it just answered from memory. The biggest problem area was understanding and comparing multiple images at once.

Why does this matter? Plenty of real-world tasks work this way. Planning a trip from a travel guide and photos, checking product details across an ad and an instruction manual, or researching a hobby using charts and diagrams all mix text with visual clues. This benchmark is a smarter way to stress-test AI before we trust it to do that kind of work. It shows there is still a long way to go before AI can truly "see" the full picture.

Key Points
  • Researchers created 1,863 real-world questions that require AI to combine information from text and images.
  • Current AI systems often fail to get the full chain of evidence, especially when they need to compare multiple images.
  • The study suggests that a complete search for relevant facts matters more than simply making the AI bigger or smarter.

Why It Matters

If AI can’t connect text and photos reliably, it can’t handle everyday research tasks for travel, purchases, or learning.

📬 Get the top 10 AI stories daily