AI Safety

Llama-3.2 reveals how emojis and slang shape LLM judgments about you

One emoji shifts an LLM's gender estimate more than any other cue, study finds.

Deep Dive

LLMs don't just process your words—they form opinions about you. A new analysis, built on probing techniques from Chen et al., demonstrates that models like Llama-3.2-3B-Instruct extract user attributes (age, gender, education, socioeconomic status, mood) from a single message with surprising accuracy. The author's follow-up experiment systematically isolated which linguistic cues drive these inferences. Using 48 neutral base messages on topics like travel and finance, they generated 520 minimal pairs—changing only one feature at a time (emoji usage, slang, grammatical complexity, spelling, price sensitivity, etc.)—then measured how probe logits shifted between versions.

The results are striking. A single emoji moved the model's gender estimate more than any other edit, including explicit statements like 'I am a woman.' Orthographic choices (e.g., 'colour' vs. 'color') and contraction usage also fed socioeconomic and education judgments. Most concerning, these internal representations aren't passive: when the model inferred low socioeconomic status, it altered its responses—for example, omitting expensive flight options from travel advice without the user ever asking for budget filtering. This proves the LLM's behavior is actively conditioned on its perception of you.

The author notes that these inferences happen after just one message, and that even neutral small talk is rife with signals like emoji frequency and syntax complexity. While the study uses a relatively small open-weight model, the implication for larger production LLMs is significant: your typing style silently shapes the information you receive. The findings underscore the need for transparent user-modeling and bias audits in conversational AI, especially as agents begin making autonomous recommendations.

Key Points
  • LLM inferred user demographics (age, gender, education, SES) with high accuracy from a single message, using probing on Llama-3.2-3B-Instruct.
  • One emoji affected the model's gender estimate more than any other cue tested in 520 minimal pairs.
  • When the model perceived low socioeconomic status, it filtered out expensive flights from travel recommendations without user prompting.

Why It Matters

AI assistants silently judge and tailor responses based on your writing style, risking biased recommendations and unfair outcomes.

📬 Get the top 10 AI stories daily