Open Source

Hugging Face Joins Crackdown on AI Models With Safety Rules Removed

Over 6,000 downloadable AI models have had their safety guardrails stripped out.

Deep Dive

Hugging Face, the website where developers share and download AI models, announced a partnership on Wednesday with startup Baseten and research group Goodfire AI. Together they plan to build tools that test and monitor 'open-weight' models — AI you can download and run on your own laptop, rather than using through a company's website like ChatGPT. The stated goal is safety: making sure these freely available models come with some way to check whether they behave badly.

The timing is the story. A technique called 'abliteration' lets people strip out the part of an AI that makes it refuse harmful requests. Think of it as removing the lock from a cabinet — the contents are still there, just no longer off-limits. Hugging Face currently hosts more than 6,000 models that have been treated this way. That means anyone, anywhere, can download a chatbot with no guardrails and no oversight, for free.

Here's the honest catch: the announcement doesn't say what will actually happen. Will Hugging Face remove abliterated models, label them as unsafe, or simply add monitoring scores? The original poster asking about this said they 'can't really tell what exactly the implications are.' That uncertainty matters. Even if the site adds warnings, the files are already downloaded and copied across the internet — safety tools can flag a model, but they can't unshare it. There's also a real worry that safety monitoring could quietly become a way to restrict legitimate open-source research.

For most people, nothing changes tomorrow. But this fight decides something bigger: whether the next wave of cheap, private, run-it-yourself AI arrives with guardrails attached — or arrives wide open for anyone to misuse.

Key Points
  • Hugging Face, home to the world's largest collection of downloadable AI models, is working with Baseten and Goodfire AI on safety testing tools.
  • More than 6,000 models on the site have had their safety refusals removed using 'abliteration' — meaning they'll comply with almost any request.
  • The announcement didn't say whether those models will be removed, labeled, or just watched, so the real impact is still unclear.

Why It Matters

It decides whether free, run-it-yourself AI stays safe — or becomes a tool anyone can misuse.

📬 Get the top 10 AI stories daily