ARC's MSP pipeline infers catastrophic AI behaviors without waiting for failures
New method estimates model safety from internal algorithms, not just test samples
ARC's proposed pipeline for aligning powerful AI, inspired by the Matching Sampling Principle (MSP), would monitor training to detect internal structure, convert that structure into advice for mechanistic estimators, then estimate rare catastrophic failure probabilities—without relying on black-box evaluation. This approach aims to flag deceptive alignment and reward hacking early in training.
- ARC's Matching Sampling Principle (MSP) enables estimating model behavior from weights without input-output samples
- The pipeline infers rare catastrophic behaviors like deceptive alignment before they ever occur in practice
- Requires a mathematically-defined goodness function and tools to convert internal structure into estimator advice
Why It Matters
Could provide a rigorous mathematical foundation for verifying AI safety before deployment, reducing reliance on empirical testing alone