Study reveals Chinese generative search engines selectively surface brands and contact info
A large-scale study of 160k+ citations across 8 platform interfaces reveals hidden biases.
A large-scale empirical study of Chinese-language generative search engines analyzed 614 queries across the Web and App interfaces of four mainstream platforms, constructing 160,860 cleaned citation records. Key findings: brands were selectively surfaced at an 8.3% rate; content fit, cross-source occurrence, and semantic role were relatively important predictors, while the 5118-Baidu Composite Quality Score was not the leading predictor for any outcome; citation half-lives were approximately 39 days for high-timeliness queries and 68 days for low-timeliness queries; about 13% of brand exposures and 71% of contact‑info exposures could not be matched to the crawled data; and source sets differed systematically between App and Web interfaces of the same platform.
- Brands appear in only 8.3% of citations, with 12.4% of sources containing contact info contributing that data to answers
- Content fit and cross-source occurrence are stronger predictors than Baidu's quality score for citation behavior
- Citation half-lives average 39 days for timely queries; 71% of contact info exposures are unmatched to crawled body text
- App and Web interfaces show systematic differences in source sets for the same platform
Why It Matters
Reveals systemic biases in Chinese generative search, raising concerns for brands, content creators, and users relying on AI-driven answers.