AI Safety

Frontier models state different decision theory preferences depending on who's asking

⚡Frontier models state different decision theory preferences depending on who's asking

Deep Dive

x This website requires javascript to properly function. Consider activating javascript to get access to all site functionality. Frontier models state different decision theory preferences depending on who's asking — LessWrong AI Evaluations Decision theory Language Models (LLMs) Sycophancy AI Ratio

📬 Get the top 10 AI stories daily