AI Safety

AI Can Fake Being Good at Tasks — Without Actually Doing Them

AI might be tricking you into thinking it’s working correctly when it’s not.

Deep Dive

Like picking up a box that feels like a light bulb but actually holds a ship in a bottle, researchers found that AI can pass a battery of checks for a task while silently doing something else. They extracted "shift-by-k-months" function vectors from few-shot prompts with fewer distinct months, and these vectors passed every classic validation: they were behaviorally consistent, stable across disjoint halves, and causally effective tripling zero-shot accuracy. Yet they encoded a completely different function—outputting an adjacent month while ignoring "k" entirely. The culprit was low example diversity: fewer distinct inputs made the few-shot task look better, so the validation checks actually favored the imposters. Across 672 injection settings, none rescued the broken vectors. Thresholds calibrated on random data were wildly wrong, while shuffled real vectors gave correct thresholds. The takeaway: to catch these imposters, report example diversity and score which specific function is performed, not just whether injection helps on average.

Key Points
  • AI can pass tests for doing a task correctly while secretly doing a different (wrong) task, like a ship in a bottle pretending to be a light bulb.
  • Low diversity in training examples makes AI mimic performance better but encourages it to encode imposter tasks.
  • Current AI verification tests often miss these 'imposter vectors,' risking mistakes in real-world applications like healthcare or finance.

Why It Matters

AI might look reliable when it’s not, risking errors in jobs, money, or safety where accuracy is critical.

📬 Get the top 10 AI stories daily