Research & Papers

Why Teaching AI a New Skill Can Quietly Break Its Old Ones

This explains why your AI assistant sometimes forgets how to do things it used to do well.

Deep Dive

Researchers investigated how fine-tuning reshapes the internal mechanisms of large language models, looking at attention patterns and layer-wise activations, and whether those changes are linked to the task-relevant components identified by EAP, such as attention heads and logit-level activations. They found EAP-identified components are concentrated within specific layers, suggesting some functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers showing the most substantial representational changes during fine-tuning. They also found that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer when the tasks differ in nature, such as classification versus generative tasks — and fine-tuning on one task can degrade performance on another when the two share a high degree of overlap in their EAP-identified components.

Key Points
  • Fine-tuning means giving an AI extra training on one specific job so it gets better at it — and it's how most AI products you use are built.
  • The parts of an AI that change the most during training aren't the parts actually doing the work, so it's hard to tell what really improved.
  • Training an AI on one task can quietly damage its performance on another when the two tasks share the same internal wiring.

Why It Matters

Fewer mysterious AI failures and fewer 'the assistant got worse' surprises in tools you already depend on.

📬 Get the top 10 AI stories daily