AI Safety

Silent Updates Scorecard: 9 AI providers fail disclosure test

Researchers found no AI provider lets outsiders verify the served model matches the evaluated one.

Deep Dive

New research from Sophia Abraham and Ben Bucknall (arXiv:2608.11803, accepted to AIES-26) exposes a critical blind spot in AI governance: silent updates. Deployed foundation models are rarely static—providers can tweak behavior via fine-tuning, classifier updates, system prompt revisions, retrieval changes, or routing changes—all without public disclosure, version bumps, or re-evaluation. This practice breaks the 'chain of custody' that current AI governance frameworks assume: that the model evaluated in a system card or benchmark is the exact artifact being served to users.

The authors audited disclosure practices across 9 first-party API providers and 7 third-party inference hosts. Their finding is stark: while providers publish extensive safety documentation and version-specific reports, none publish enough information for an external party to verify the served artifact matches the documentation. To address this, they introduce the Silent Updates Scorecard, a public instrument for measuring post-deployment disclosure transparency. They also propose a Three-Part Behavioral Trigger System that defines when modifications to a deployed system should trigger a disclosure or re-evaluation obligation. For enterprises and regulators relying on model audits, this paper signals that current verification methods are fundamentally insufficient—and offers concrete tools to start closing the gap.

Key Points
  • Studied 9 first-party AI API providers and 7 third-party inference hosts; none allow external verification of served model vs. documented model.
  • Identified silent update vectors: fine-tuning, classifier updates, system prompt revisions, retrieval changes, routing changes.
  • Proposes Silent Updates Scorecard and Three-Part Behavioral Trigger System to force disclosure/re-evaluation when behavior changes post-deployment.

Why It Matters

Breaks the assumption of model verifiability—regulatory audits and enterprise evaluations can no longer trust a system card without external checks.

📬 Get the top 10 AI stories daily