MIRAGE framework boosts MSR dataset discoverability and reuse
New metadata enrichment method using LDA and Semantic Scholar API
A new paper from researchers Aabia Ather, Muhammad Usayd Ather, Qurat-Ul-Ain Somroo, and Muhammad Khuram Shahzad introduces MIRAGE (Metadata-Integrated Repository Analysis and Guided Enhancement), a framework designed to improve the analysis and discoverability of Mining Software Repositories (MSR) datasets. The study expands on previous dataset directories by enriching metadata categories, adding advanced filtering options, and performing FAIRness assessments. Using the Semantic Scholar API, the team collected metadata from MSR papers published between 2013 and 2024, applying Latent Dirichlet Allocation (LDA) for topic modeling alongside statistical analysis.
The enhanced directory includes dataset-level attributes such as repository hosting site, data format, accessibility, reusability, and quality. Key findings reveal that the choice of repository hosting site and data format significantly impacts citation patterns and dataset usability. By improving metadata quality and enabling topic-driven analysis, MIRAGE supports more effective reuse and evaluation of research artifacts. The framework offers automated metadata enrichment and enhanced filtering, making it easier for researchers to locate and assess relevant datasets for empirical software engineering studies.
- Metadata enriched for over a decade of MSR papers (2013–2024) via Semantic Scholar API
- LDA topic modeling and statistical analysis reveal hosting site and format influence citation patterns
- New directory includes dataset-level attributes: hosting site, format, accessibility, reusability, quality
Why It Matters
Better metadata and FAIRness assessment make software repository datasets more discoverable and reusable for researchers.