Google’s New LLMs Automate Biomedical Screening With 74% Accuracy — Here’s Why Researchers Are Paying Attention
Gemini 2.5 Pro and Gemma 3 models team up to automate EQ-5D screening in PubMed abstracts
Researchers developed an ensemble of Google's LLMs—Gemini 2.5 Pro, Gemma 3-12B, and Gemma 3-27B—to automate EQ-5D study detection in PubMed abstracts. Their weighted ensemble achieved a 0.74 F1-score and 0.74 accuracy, outperforming individual models and improving balance between precision and recall. The soft-stacking meta-classifier added reliability and interpretability for systematic literature reviews.
- Gemini 2.5 Pro + Gemma 3-12B + Gemma 3-27B ensemble reached 0.74 F1-score and 0.74 accuracy on EQ-5D detection in PubMed abstracts
- Soft stacking meta-classifier improved reliability and interpretability over individual models
- Automated screening reduces manual review workload by 74% for systematic literature reviews
Why It Matters
This ensemble approach could cut weeks off biomedical literature reviews, accelerating evidence synthesis and clinical guideline development.