Bias & Persona

Understanding the biases language models absorb and the personas they adopt, and how both shape model behavior.

The Problem

Language models inherit biases from their training data, and their behavior can shift depending on the persona they adopt, sometimes reinforcing harmful stereotypes or producing inconsistent outputs. These effects are subtle and context-dependent, making them hard to measure and hard to fix without eroding a model's usefulness.

What We're Working On

  • Bias measurement — Detecting and quantifying demographic and social biases in model outputs.
  • Persona analysis — How models represent and switch between personas internally, and the effect on safety.
  • Mechanistic links — Connecting bias and persona to internal representations, so they can be steered rather than patched.

Related Publications

1 paper in Bias & Persona

View all