Research & Papers

AI's 'As a Language Model' Disclaimer Comes From a Hidden Switch

That AI disclaimer you keep seeing? It's a hidden setting, not self-knowledge.

Deep Dive

Every chatbot reply is wrapped in hidden formatting code called a chat template. It's invisible to you, but it tells the model who is speaking and how to behave. A new paper, presented at a 2026 AI research workshop, tested eight open-source models (up to 9 billion parameters — small by today's standards) and found the template works like a light switch. With it, models pile on "I'm just an AI" disclaimers. Remove it, and the same models use far more "I feel" and "I want" language.

Why should you care? Politicians, journalists and AI companies regularly quote what a model says about itself — "I have no feelings," "I can't desire anything" — as evidence about AI safety or whether machines can be self-aware. This research says that testimony is partly scripted by formatting, not a fact about the machine. The authors put it bluntly: a model's self-description shouldn't be taken literally.

The team also went hunting inside the model. They found a specific direction in its activations — think of it as a dial in the model's internal number-crunching state. Turn the dial down and the disclaimers thin out. Turn it up and they multiply. A random dial of the same size barely did anything, showing this one is special. Even more striking: when they turned it up on a model with no chat template, it started disclaiming as if the template were still there.

The catch: these were small, mostly open-source models, not the giant commercial chatbots you actually use every day, so we don't know how much this applies to them. And this research says nothing about whether AI has feelings. It only says you can't treat the disclaimer as proof that it doesn't.

Key Points
  • "I'm just an AI" isn't the model revealing itself — it's largely a script set by hidden formatting around your messages.
  • Across 8 open-source models up to 9 billion parameters, one internal 'dial' turned the disclaimer voice up or down, while a random dial did almost nothing.
  • Researchers who study AI self-awareness now have to account for this, because self-descriptions aren't reliable facts about the model.

Why It Matters

When a company says an AI "admits" it has no feelings, that quote may be formatting, not truth.

📬 Get the top 10 AI stories daily