ACM RecSys reproducibility survey: 51 papers show gap between ideal and practice
Six years of RecSys reproducibility papers reveal a surprising lack of methodological diversity.
A new survey of the ACM RecSys Reproducibility Track, covering 51 accepted papers from 2020 through 2025, provides the first structured look at how recommender system researchers actually operationalize reproducibility. Authors Alan Said and Alejandro Bellogin classify contributions by type and analyze common patterns in datasets, algorithms, frameworks, and evaluation practices. They find that the track has significantly expanded in scope since its 2020 launch: what began as a focus on reproduction and replication now includes benchmarking, resource contributions, and methodological studies. However, the work consistently relies on a narrow palette — a limited set of benchmark datasets, well-known algorithms, and standard evaluation protocols.
Perhaps most revealing, the survey finds that reproducibility in practice often means extending prior experiments rather than strictly replicating them. Studies frequently introduce additional models or evaluation criteria, blurring the line between reproduction and novel contribution. While these efforts have undeniably improved transparency and documentation within the RecSys community, they have had limited impact on methodological diversity. The authors conclude that a gap persists between the conceptual definition of reproducibility and how it is actually implemented, suggesting the community must rethink incentives and standards to close that divide.
- Analyzed 51 papers from the ACM RecSys Reproducibility Track (2020-2025) across reproduction, replication, benchmarking, and resource contributions.
- Reproducibility studies rely on a limited set of datasets, algorithms, and evaluation protocols, rarely introducing methodological diversity.
- Most studies extend prior experiments (adding models or criteria) rather than strictly replicating them, creating a gap between conceptual reproducibility and practice.
Why It Matters
Recommender system developers must prioritize methodological diversity, not just replication, to ensure reliable and innovative results.