New Science Quiz Stumps Top AI Chatbots — Here's Why That Matters
Researchers built a 500-question biology exam, and the best AI barely passed a third of it.
A researcher has released a new exam for AI, called BioPhys-Bridge, designed to test something most chatbots are quietly bad at: reading real science. The 500 cases come from biophysics, a field where physics and biology overlap — think how proteins fold, how cells sense forces, or how electricity moves through nerves. To answer correctly, an AI has to find the right evidence in a paper, plug numbers into a physics equation, and then explain what it means for a living organism. That's three hard steps, not one.
Why should you care? Because AI chatbots are increasingly used to summarize medical studies, suggest treatments, and help design experiments. If an AI skips a step or invents a number, the answer can sound perfectly confident and still be wrong. That's called hallucination — the AI making things up. In this test, the strongest model, DeepSeek-V4-Flash, scored 0.36 on matching the correct evidence. In plain terms: it cited the right source roughly one time out of three. Qwen3.7-Max and GPT-4o-mini scored lower still.
The test also uses something called RAG, which just means letting AI look things up in a document instead of relying on memory. Even with the source material sitting right in front of them, the models struggled to point to the exact right passage. The benchmark itself covers six biological topics and nine types of physics models, and only 81 of the 500 cases were double-checked by human experts — so this is an early, imperfect measuring stick, not the final word.
The honest catch: this is a test, not a product. It doesn't make AI smarter, and scores will likely climb as models improve. But the takeaway for you is simple. AI is genuinely useful for explaining concepts and saving reading time. It is still shaky at precisely citing evidence in complex science. So if a chatbot gives you a medical or scientific claim, ask for the source — then check it yourself.
- A new 500-question test checks whether AI can read real biophysics papers without inventing facts.
- The best model scored 0.36 on pointing to the correct evidence — right about one time in three.
- AI is fine for explaining ideas, but you should still verify sources on health and science claims.
Why It Matters
If you use AI for health or science questions, this shows why double-checking its sources still matters.