AI Safety

Bluesky's AI Moderators Miss Most Harmful Posts, Study Finds

The system flags bad content in seconds — but lets roughly 4 in 5 harmful posts through.

Deep Dive

A team of researchers just completed the first large-scale outside audit of how Bluesky — the social network that grew fast after people left X — polices its own content. They could do this because Bluesky publishes its moderation records publicly, which Facebook, Instagram and X do not. The study examined 10.6 million moderation labels applied in 2025, essentially a full year of the platform's decisions.

The picture that emerges is a partnership between machines and people. Labels for sexual and graphic content are applied automatically, often within seconds. But tougher calls — things like harassment or coordinated hostility toward particular groups — get pushed to human reviewers and can take hours or even days to resolve. In practical terms, that means the most disturbing content may also be the slowest to disappear, leaving a window where it stays visible.

The accuracy findings are the most striking. When the system flags a post, it's right about 84 percent of the time, which researchers call high precision. But it only catches about 22 percent of harmful content — the rest slips past. In a random sample, human reviewers found 4.5 times more harmful material than the automated system did. In plain terms: Bluesky is good at not making false accusations, and bad at finding everything it should.

The takeaway for everyday users is simple. A platform saying it has moderation does not mean the worst content is gone, so your own filters, mutes and blocks still do real work. The bigger story is transparency: because Bluesky's logs are open, outsiders could finally measure what AI moderation actually does. That same scrutiny has been impossible on larger platforms, and this study is a preview of what we might learn if that changed.

Key Points
  • Bluesky's automated system catches sexual and graphic content in seconds, but slower human review is needed for harassment and other complicated cases, which can take days.
  • The system catches only about 1 in 5 harmful posts, though it is accurate about 84 percent of the time when it does flag something.
  • The study was only possible because Bluesky makes its moderation records public — Facebook, Instagram and X do not.

Why It Matters

Social apps lean on AI to police content; this shows how much still slips through, so your own filters matter.

📬 Get the top 10 AI stories daily