Detecting and reducing scheming in AI models
OpenAI and Apollo Research have developed new evaluations to detect 'scheming' or hidden misalignment in frontier models. They've also shared an early method to reduce these behaviors, providing critical stress tests for AI safety.