Superalignment
Researching how increasingly capable AI systems can remain aligned with human intent.
Superalignment is a restricted Mecha ML research program focused on alignment methods for highly capable AI systems. The work explores scalable oversight, evaluation, interpretability, robustness, and control techniques intended to help keep advanced systems understandable and aligned with human intent.
Key Features
Scalable Oversight
Study supervision methods that can remain useful as model capabilities increase.
Alignment Evaluations
Develop evaluations for robustness, deceptive behavior, goal misgeneralization, and control failures.
Interpretability & Control
Explore tools for understanding model behavior and constraining high-risk actions.