Research & Papers

Google’s New LLMs Automate Biomedical Screening With 74% Accuracy — Here’s Why Researchers Are Paying Attention

Gemini 2.5 Pro and Gemma 3 models team up to automate EQ-5D screening in PubMed abstracts

Deep Dive

Researchers developed an ensemble of Google's LLMs—Gemini 2.5 Pro, Gemma 3-12B, and Gemma 3-27B—to automate EQ-5D study detection in PubMed abstracts. Their weighted ensemble achieved a 0.74 F1-score and 0.74 accuracy, outperforming individual models and improving balance between precision and recall. The soft-stacking meta-classifier added reliability and interpretability for systematic literature reviews.

Key Points
  • Gemini 2.5 Pro + Gemma 3-12B + Gemma 3-27B ensemble reached 0.74 F1-score and 0.74 accuracy on EQ-5D detection in PubMed abstracts
  • Soft stacking meta-classifier improved reliability and interpretability over individual models
  • Automated screening reduces manual review workload by 74% for systematic literature reviews

Why It Matters

This ensemble approach could cut weeks off biomedical literature reviews, accelerating evidence synthesis and clinical guideline development.

📬 Get the top 10 AI stories daily