Research & Papers

PERCEPT: First Persian-English code-mixed corpus with 6,800 UD-tagged posts

New corpus from 6,800 X, Instagram, and Digikala posts unlocks code-mixing analysis

Deep Dive

Code-mixing—blending multiple languages in one utterance—is common on social media, but Persian-English data has been largely missing. To fix this, a team led by Ghazal Kalhor and Behnam Bahrak built PERCEPT, the first publicly available large-scale corpus of this type. The dataset pulls 6,800 posts from X, Instagram, and Digikala, and applies Universal Dependencies (UD) part-of-speech tags to every code-mixed word. Using an LLM-assisted annotation pipeline, the authors automatically assigned POS tags and document-level topics; human evaluation confirmed high agreement with gold annotations, validating the approach.

Analyzing PERCEPT, the researchers found that nouns are the most common code-mixed category across all platforms, but the distribution of other POS tags varies. Interestingly, positional patterns of code-mixed words were remarkably consistent, while the “triggering effect”—where one code-mixed word increases the chance of another—was substantially stronger in Digikala comments. PERCEPT is publicly available, giving NLP researchers a valuable resource for syntax-aware models, multilingual understanding, and Persian-English code-switching studies.

Key Points
  • 6,800 posts collected from X, Instagram, and Digikala
  • First Persian-English code-mixed corpus with Universal Dependencies POS tags
  • LLM-assisted annotation shows high agreement with human gold annotations
  • Nouns are the dominant code-mixed POS category, with platform-specific variations

Why It Matters

Gives NLP teams the first benchmark for Persian-English code-mixing, enabling syntax-aware multilingual models and deeper linguistic insights.

📬 Get the top 10 AI stories daily