AI Safety

Researchers show AI models can hide covert messages in latent space

A new attack relocates cloud clusters in spiking networks, bypassing 9.7:1 shift ratios to evade detection.

Deep Dive

Igor Pereverzev explores how AI language models like SpikeGPT might evade a monitor by relocating message clouds in latent space instead of scrambling them. Using a 216M parameter spiking RWKV model, the post initially cited a ratio of shift to deformation around 9.7 as evidence of relocation—but then admits that number fooled him and doesn't prove the classes moved rather than collapsed. The work focuses on the geometry of representations rather than the architecture, with implications for security and neurointerfaces.

Key Points
  • SpikeGPT (216M params, 5B tokens of OpenWebText) used to demonstrate covert channel attacks via latent space relocation
  • Attack shifts message cloud centroids by 2.5 units while keeping spread changes at just 0.26, evading detection
  • Spiking neural networks (SNNs) enable sparse, event-driven computation where spikes (0/1) carry information via timing

Why It Matters

Demonstrates fundamental vulnerabilities in AI models' internal representations, threatening security and enabling hidden communication channels in deployed systems.

📬 Get the top 10 AI stories daily