LIWC's limited predictive role in depression classification revealed
New controlled substitution study across 5 corpora challenges LIWC's predictive utility
Researchers from multiple institutions (including Boyd and Sisman) conducted a rigorous evaluation of Linguistic Inquiry and Word Count (LIWC) for depression-related classification. They compared intact LIWC features against three controlled substitutes: a PCA-rotated version that removed direct access to category coordinates, a participant-shuffled version that broke person-level alignment, and a random-marginal version that preserved feature distributions. The evaluation spanned five English and Chinese depression-related speech/text corpora under matched participant-level cross-validation. The primary question: Does LIWC provide incremental predictive value beyond the raw language features themselves?
The results showed limited evidence for stable LIWC gains under frozen, participant-level early fusion. None of the pre-specified dataset-blocked contrasts survived multiple-comparison correction, indicating that any observed improvements were not statistically reliable. A separate calibration with SBERT confirmed the procedure could detect larger signals when present, but the five-corpus power limitation remained. The authors conclude that LIWC's value lies in its auditable, corpus-conditioned interpretive capabilities rather than predictive performance. Importantly, these findings do not apply to fine-tuned, sequence-aware, or end-to-end architectures, leaving open the possibility that LIWC could benefit more advanced models.
- Evaluated LIWC across 5 English and Chinese depression-related corpora using three controlled substitution methods (PCA-rotated, shuffled, random-marginal)
- No significant predictive gains after multiple-comparison correction in frozen early fusion settings
- LIWC remains useful as an auditable interpretive layer but not as a performance booster in these architectures
Why It Matters
Caution against assuming linguistic categories directly improve depression detection; interpretability and prediction are separate goals.