Research Areas
Exploring critical sub-sections of AI alignment and safety research.
Alignment & Coordination
Ensuring AI systems (individually and in groups) reliably understand and act on human intentions.
Explore research & publications
Bias & Persona
Understanding the biases language models absorb and the personas they adopt, and how both shape model behavior.
Explore research & publications
Governance
Studying the policy, legal, and institutional frameworks needed to govern AI safely and responsibly.
Explore research & publications
Interpretability
Reverse-engineering the internal mechanisms of neural networks to predict, verify, and steer their behavior.
Explore research & publications
Metacognition & Hallucination
Investigating whether models know what they know, and why they confidently state things that are false.
Explore research & publications
Security & Robustness
Probing and hardening AI systems against adversaries — from prompt injection and red-teaming to adversarial robustness.
Explore research & publications