Research & Papers

Zhang et al. build multimodal dataset boosting keyword extraction with images and audio

1000 academic papers now include audio and visual text for richer keyword extraction.

Deep Dive

Traditional keyword extraction methods rely exclusively on textual data, ignoring visual details from images and audio features from speeches or presentations within academic papers. To address this gap, Jingyu Zhang and four co-authors created a multimodal dataset of 1,000 academic papers, each sample including the paper's written text, extracted text from embedded images, transcribed audio content, and ground-truth keywords. This resource aims to enrich information diversity and capture cross-modal correlations that improve representation learning.

Using both unsupervised and supervised keyword extraction models, the team ran experiments with each modality in isolation and with fused text from all three. Results show that text from different modalities exhibits distinct characteristics in the model, and concatenating paper text, image text, and audio text effectively boosts keyword extraction performance over any single modality. The dataset and findings provide a foundation for future multimodal NLP research in academic information retrieval.

Key Points
  • Dataset includes 1,000 academic paper samples with text, images, audio, and keywords.
  • Existing keyword extraction relies solely on text; this work adds visual and audio modalities.
  • Concatenating all three textual modalities outperforms using paper text alone.

Why It Matters

Enables more accurate keyword extraction for research discovery by leveraging image and audio signals from papers.

📬 Get the top 10 AI stories daily