Eigenbraid's Glasswing Standard proposes phased AI safety transparency and release controls
A concrete first-step policy for AI safety based on Anthropic's Project Glasswing approach.
Eigenbraid's Glasswing Standard, inspired by Anthropic's Project Glasswing, outlines a phased approach to AI safety that prioritizes transparency before control. Phase 1 creates a standardized early-access group comprising major labs and government bodies. These organizations issue regular public reports on whether a model is safe, with prediction markets extrapolating release dates based on accumulating sign-offs. Early access users are purely advisory, building a track record of forecasting strengths and weaknesses without liability. The proposal notes that this system already paid off: early access to the Mythos model revealed a concrete, graphable spike in cybersecurity capabilities, raising awareness without causing harm.
Phase 2 introduces formalization as capabilities grow more dangerous. Labs may initially retain full veto power, but oversight can tighten through super-majority voting requirements among early-access users or formalized red-team evaluations using the same models. The goal is to eventually reduce bug reports and dangerous capabilities to a trickle before public release. Eigenbraid emphasizes aligning incentives—labs profit from releases but face a mild speed-bump that incentivizes safeguards, while the government balances economic and security interests. This framework avoids slowing frontier development, builds public consensus on ASI threats, and provides a scalable roadmap for future safety measures.
- Phase 1 transparency: standardized early-access group issues public reports, enabling prediction markets and track-record building without liability.
- Phase 2 formalization: super-majority approval and red-team evaluations tighten oversight based on capability evolution.
- Early access to Mythos model demonstrated value by revealing a cybersecurity capabilities spike before public release.
Why It Matters
A pragmatic, low-friction framework to build AI safety track records without stalling innovation.