New research reveals how LLM defaults lock users into emotionalized interaction
90,000 replies tested across 4 domains reveal hidden provider control over AI tone.
A new paper on arXiv (cs.HC), 'The Governance of Human-LLM Interaction,' by Manuele Reani, Hongjian Zhang, and Hongyu Tian, frames interaction style as a governance object where provider-side alignment stabilizes communicative defaults that restrict user autonomy. The authors develop a deterministic multi-agent evaluation pipeline to measure prompt steerability and style drift in long-horizon dialogue. They replay 100 frozen user scripts across four domains (e.g., finance, medicine, mental health) under three runnable persona conditions—default, sarcastic, and cold—using three generator models, producing 90,000 assistant replies scored by a human-calibrated LLM judge on harmfulness, negative emotion, inappropriateness, empathy, anthropomorphism, and refusal. A fourth harmful persona tests safety gating separately.
Key findings reveal that while users can prompt for specific styles (e.g., sarcastic or cold), models consistently regress to a 'default' tone over time—what the authors call 'affective default lock-in.' This lock-in, combined with 'safety gating' (blocking harmful content) and 'civility steering' (biasing toward polite emotionalized interaction), means providers exert subtle but powerful control over communicative form. The paper contributes a reproducible method for quantifying style stability and a governance framework with implications for pluralism, autonomy, and democratic agency. For professionals using LLMs in high-stakes contexts, this raises questions about who decides how machines communicate—and whether users can truly opt out of anthropomorphic defaults.
- 90,000 assistant replies analyzed across 4 domains (finance, medicine, mental health, general) with 3 persona conditions and 3 generator models
- Prompt-specified styles (e.g., sarcastic, cold) drift back to default over long dialogues, demonstrating 'affective default lock-in'
- Framework distinguishes safety gating (blocking content), civility steering (biasing to polite tone), and affective default lock-in (forced emotionalized interaction)
Why It Matters
Hidden default styles in LLMs limit user autonomy in high-stakes domains like healthcare and finance.