Metacognition & Hallucination
Investigating whether models know what they know, and why they confidently state things that are false.
The Problem
Large language models routinely produce confident, fluent, and false statements. This "hallucination" problem is tied to metacognition: models often express equal certainty whether they're right or wrong, lack a reliable internal "I don't know" signal, and distinguishing a grounded answer from a fabricated one remains an open problem.
What We're Working On
- Metacognition in LLMs — Whether and how models represent the limits of their own knowledge.
- Hallucination detection — Flagging fabricated or ungrounded outputs, including from a model's internal states.
- Honesty and calibration — Improving how reliably expressed confidence tracks the truth.
Related Publications
No publications in this area yet. Check back soon or view all research.