Metacognition & Hallucination

Investigating whether models know what they know, and why they confidently state things that are false.

The Problem

Large language models routinely produce confident, fluent, and false statements. This "hallucination" problem is tied to metacognition: models often express equal certainty whether they're right or wrong, lack a reliable internal "I don't know" signal, and distinguishing a grounded answer from a fabricated one remains an open problem.

What We're Working On

  • Metacognition in LLMs — Whether and how models represent the limits of their own knowledge.
  • Hallucination detection — Flagging fabricated or ungrounded outputs, including from a model's internal states.
  • Honesty and calibration — Improving how reliably expressed confidence tracks the truth.

Related Publications

No publications in this area yet. Check back soon or view all research.