AI Safety

Proposal for tracking the effects of architecture on monitorability

Proposal for tracking the effects of architecture on monitorability

Deep Dive

Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward). [1] As companies begin to explore such a

📬 Get the top 10 AI stories daily