Research & Papers

Researchers Clone AI Models by Asking 8,000 Smart Questions

Your company's secret AI could be copied just by asking it the right questions.

Deep Dive

Think of a top-secret recipe. You can't see the ingredients, but by tasting the dish in a thousand tiny variations, you can work out the recipe. That's what this new research does to AI models. The paper shows that transformer models — the kind behind ChatGPT and image recognition — leak their internal structure when you ask them questions that measure how their outputs bend and twist. The researchers call this "curvature cryptanalysis" because it uses the shape of the AI's responses to break into its hidden design.

In practice, the attack is shockingly cheap. The team needed only 8,193 questions to extract the "hidden directions" inside the model — the internal patterns that guide how it thinks. Those directions were recovered with more than 94% accuracy. Once they had those patterns, they could build a working copy of the AI that agreed with the original over 93% of the time, while staying almost as accurate on standard tests. It's like photocopying a brain, but instead of reading the neurons, you just shine light through it and trace the shadows.

This matters because AI models are expensive to build. Companies guard them like trade secrets, and many people assume that keeping the code secret is enough. This research shows that secrecy isn't enough. If you can talk to an AI, you can potentially steal it. The researchers also found that adding noise or rounding the outputs slows the attack down, but a determined attacker can adapt and still succeed. It's not a full fix, just a speed bump.

The good news? This kind of work also helps researchers understand AI better. If you can see how a model stores knowledge, you might find ways to make it more transparent or fix its biases. But for anyone running a paid AI service, this is a warning: your model's behavior is a fingerprint that can be copied. The question is how to protect it without hiding it away entirely.

Key Points
  • A new technique copies AI models by analyzing how their outputs bend and curve, using only 8,193 questions.
  • The copied models match the original's decisions over 93% of the time, even though the attacker never saw the original code or internal settings.
  • Adding noise or rounding outputs slows the attack but doesn't stop it — a determined attacker can adapt and still succeed.

Why It Matters

Secrets inside AI models aren't safe — any public AI can be reverse-engineered from its answers alone.

📬 Get the top 10 AI stories daily