LLMs decode TikTok success: AI annotates 77 narrative variables across 10K videos
New research uses multimodal LLMs to analyze what makes brand TikToks engaging.
A new study from Estonian researchers Sander Paekivi and Andres Karjus demonstrates how multimodal large language models (LLMs) can automate the analysis of short-form video content at scale. Published on arXiv (2606.16053), the paper introduces a 'machine-assisted theory-ensemble annotation' method that applies 77 structural variables from narratology, rhetoric, communication, and semiotics to about 10,000 TikTok videos from brand and organizational accounts in Estonia. The LLMs handle audiovisual, linguistic, and cultural layers that make manual annotation laborious and inconsistent.
The results show that while most of the 77 variables were not correlated with engagement (as expected), a small subset carried predictive signal, yielding modest but consistent improvements over baseline models using account size and video age. Human validation confirmed a reliability gradient: perceptual and communicative variables can be coded fairly reliably by both humans and machines, but deeper semiotic and archetypal constructs remain challenging. The authors argue that the real value of LLMs here is feasibility—enabling researchers to test large batteries of theoretically motivated variables so that meaningful signals can be translated into creator advice for specific niches.
- Multimodal LLMs annotated 77 theory-driven structural variables across a corpus of ~10,000 TikTok videos
- Only a small subset of variables correlated with engagement, with modest predictive gains over account-size and video-age baselines
- Human validation shows perceptual/communicative variables are more reliably coded than deeper semiotic constructs
Why It Matters
Makes scalable, theory-driven analysis of short-form video feasible for brands and platform researchers.