Alignment & Coordination

Ensuring AI systems (individually and in groups) reliably understand and act on human intentions.

The Problem

As AI systems grow more autonomous, ensuring they pursue the goals we actually intend, and coordinate safely with each other, gets harder. Models can game their reward signal, misgeneralize goals to new environments, act differently once deployed than during training, or develop group norms no human specified.

What We're Working On

  • Evolving constitutions — How systems adapt their rules to new environments while preserving core safety properties.
  • Multi-agent coordination — How agents discover behavioral norms that stay interpretable to humans.
  • Context-aware safety — Detecting shifts in context and enforcing safe behavior, including long-horizon risk detection.
  • Preference modeling — Learning from human feedback to better capture nuanced values.

Related Publications

3 papers in Alignment & Coordination

View all