Assertive language shifts Llama-3.2 reasoning on animal welfare
How you write training data can nudge LLMs toward or away from pro-animal stances.
Researchers Jasmine Brazilek and Harper Dunn from an undisclosed institution tested how ten linguistic features in fine-tuning data change Llama-3.2-1B's preference for pro-animal-welfare reasoning. They used vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark. Eight of the ten features produced statistically significant shifts. Seven features moved the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two features—hedged language and concrete sensory description—moved the model away from a pro-welfare stance, effectively diluting it. First-person perspective had no statistically significant effect.
The practical takeaway is clear: when writing animal-welfare content that might end up in LLM training corpora, writers should assert a position rather than neutrally describe a scene. Features that shift the model are those that make the writer's stance explicit, while features that dilute it contain animal-welfare content but withhold that stance. This finding has implications for advocacy groups, content creators, and AI ethics researchers concerned with how training data shapes model behavior on sensitive topics.
- 8 of 10 linguistic features tested on Llama-3.2-1B significantly shifted its animal-welfare reasoning.
- Assertive certainty, moral vocabulary, emotion words, and narrative structure each pushed the model toward stronger pro-welfare reasoning.
- Hedged language and concrete sensory description diluted pro-welfare stance—recommendation: assert rather than describe neutrally.
Why It Matters
Shows that training data style can inadvertently bias LLMs on ethical topics like animal welfare.