Bias & Persona
Understanding the biases language models absorb and the personas they adopt, and how both shape model behavior.
The Problem
Language models are trained on human data, and they inherit both its biases and its many voices. The result is a system whose behavior can shift depending on the "persona" it adopts, sometimes in ways that reinforce harmful stereotypes or produce inconsistent, unreliable outputs.
Key challenges include:
- Inherited bias — Models absorb social, cultural, and demographic biases from training data and can amplify them at scale.
- Persona drift — A model's tone, values, and reliability can change as it takes on different roles or characters, making behavior hard to predict.
- Measurement — Bias and persona effects are subtle and context-dependent, making them difficult to detect and quantify.
- Steering without erasing — Mitigating harmful bias while preserving a model's usefulness and expressive range is a delicate balance.
What We're Working On
- Bias measurement — Developing methods to detect and quantify demographic and social biases in model outputs.
- Persona analysis — Studying how models represent and switch between personas internally, and how this affects safety and consistency.
- Mechanistic links — Connecting bias and persona behaviors to specific internal representations, so they can be understood and steered rather than merely patched.
Related Publications
1 paper in Bias & Persona