Researchers warn TikTok data tools lack reliability
TikTok's top 3 data tools give wildly different results, threatening research validity.
A team of researchers led by Gayoung Jeon from the University of North Carolina at Chapel Hill has published a peer-reviewed study exposing major inconsistencies in TikTok data acquisition tools. The paper, titled *WhichTok? Comparing Three TikTok Data Acquisition Tools*, evaluates the official TikTok Research API, Pyktok, and Apify across five endpoints: User, Hashtag, Keyword, Comment, and Related Video.
The study found that only the User endpoint produced reliable and consistent results across all three tools. However, hashtag and keyword searches yielded significantly different datasets due to methodological variations—such as the Research API using back-end API calls while Pyktok and Apify rely on front-end web scraping. These discrepancies introduce systemic biases that could compromise the validity of social science and political discourse research on TikTok. The paper, accepted for ICWSM'27, calls for standardized methodologies and ethical guidelines to address these reproducibility challenges.
- Only TikTok’s User endpoint produced consistent results across the Research API, Pyktok, and Apify tools.
- Hashtag and keyword searches showed substantial discrepancies due to differing scraping methods.
- The study highlights a reproducibility crisis in TikTok-related research, with systemic biases affecting data validity.
Why It Matters
Researchers risk flawed conclusions on TikTok’s influence without standardized data tools.