Research & Papers

OncoTriad-QA: 86K-question benchmark tests AI's pan-cancer multimodal reasoning

86.1K questions across 9,281 patients and 32 cancer types push AI beyond single-modal diagnosis.

Deep Dive

OncoTriad-QA is a new patient-level benchmark for pan-cancer question answering that forces AI models to integrate radiology, pathology, genomics, and clinical metadata—something most existing medical LLM and VLM benchmarks ignore. Created by researchers including Mubarak Shah's group, it contains 86.1K semantic questions derived from 9,281 TCGA patients across 32 cancer cohorts, aligning CT/MRI imaging, whole-slide histopathology, somatic mutations, copy-number alterations, DNA methylation, bulk RNA-seq, and clinical notes. To build the dataset, the team used a source-grounded LLM-assisted pipeline that anchors answers in curated labels, diagnostic reports, and molecular profiles, with automated consistency checks and clinician review.

Alongside the benchmark, the authors introduce OncoVLM, a multimodal model that uses learned projectors to map native radiology, pathology, methylation, and RNA-seq evidence into an LLM interface. Experiments show that both general-purpose and medical LLMs remain weak on comprehensive pan-cancer QA, particularly when questions require combining imaging findings, tumor morphology, and molecular profiles. After fine-tuning on OncoTriad-QA, OncoVLM outperformed MedGemma-4B by an average of 10.7 points across multiple-choice and open-ended questions, using both MCQ accuracy and BERTScore-F1, with consistent gains in radiology-only, pathology-only, and all-available settings. The benchmark also exposed gaps in current medical models, suggesting that patient-level reasoning across modalities is a distinct skill that needs dedicated supervision.

Key Points
  • OncoTriad-QA includes 86.1K semantic questions across 9,281 TCGA patients and 32 cancer cohorts
  • OncoVLM integrates CT/MRI, histopathology, somatic mutations, DNA methylation, RNA-seq, and clinical metadata via learned projectors
  • Fine-tuned OncoVLM beats MedGemma-4B by an average of 10.7 points on both multiple-choice and open-ended QA

Why It Matters

For precision oncology, AI must reason like a tumor board—across imaging, pathology, and genomics—and this benchmark finally measures that.

📬 Get the top 10 AI stories daily