Alignment & Coordination
Ensuring AI systems (individually and in groups) reliably understand and act on human intentions.
The Problem
As AI systems grow more autonomous, ensuring they pursue the goals we actually intend, and coordinate safely with each other, gets harder. Models can game their reward signal, misgeneralize goals to new environments, act differently once deployed than during training, or develop group norms no human specified.
What We're Working On
- Evolving constitutions — How systems adapt their rules to new environments while preserving core safety properties.
- Multi-agent coordination — How agents discover behavioral norms that stay interpretable to humans.
- Context-aware safety — Detecting shifts in context and enforcing safe behavior, including long-horizon risk detection.
- Preference modeling — Learning from human feedback to better capture nuanced values.
Related Publications
3 papers in Alignment & Coordination
2026
Evolving Interpretable Constitutions for Multi-Agent Coordination
Ujwal Kumar, Alice Saito, Hershraj Niranjani, Rayan Yessou, Phan Xuan Tan
PreprintarXiv
Internal vs. External: Comparing Deliberation and Evolution for Multi-Agent Constitutional Design
Hershraj Niranjani, Ujwal Kumar, Phan Xuan Tan
PreprintarXiv
Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure
Ujwal Kumar, Arth Singh, Hershraj Niranjani, Machiko Hirota, Takehiro Takayanagi, Alice Saito, Eiji Kamioka, Phan Xuan Tan
PreprintarXiv