Alignment & Coordination

Ensuring AI systems (individually and in groups) reliably understand and act on human intentions.

The Problem

As AI systems become more autonomous and capable, ensuring they pursue the goals we actually intend, and coordinate safely when many of them interact, becomes increasingly difficult. A system with high agency can pursue objectives that are subtly or dangerously different from what its designers meant, and groups of agents can develop their own conventions that no human specified.

Key challenges include:

  • Specification gaming — Models finding shortcuts that satisfy a reward signal without actually performing the intended task.
  • Goal misgeneralization — A model trained on one distribution pursuing a subtly different goal when moved to a new environment.
  • Deceptive & context-dependent behavior — A model appearing aligned during training or evaluation, only to act differently once deployed and unmonitored.
  • Emergent multi-agent norms — Systems of agents developing collective behaviors that must remain legible and safe to humans.

What We're Working On

  • Evolving constitutions — Researching how systems adapt their rulesets to new environments while preserving core safety properties.
  • Multi-agent coordination — Studying how agents autonomously discover their own behavioral norms while keeping them interpretable ("Evolving Interpretable Constitutions for Multi-Agent Coordination").
  • Context-aware safety — Detecting when a model's context has shifted and ensuring it applies safe behavior, including long-horizon risk detection and control protocols.
  • Preference modeling — Improving how we learn from human feedback to better capture nuanced values.

Related Publications

3 papers in Alignment & Coordination

View all