Security & Robustness

Probing and hardening AI systems against adversaries — from prompt injection and red-teaming to adversarial robustness.

The Problem

As AI is deployed into critical infrastructure and consumer products, it becomes a new, poorly understood attack surface, and a brittle one: small, deliberate input manipulations can make state-of-the-art models fail. Prompt injection can hijack a model's behavior through any text it reads, and surfacing failure modes before adversaries do requires systematic, automated red-teaming rather than ad-hoc testing.

What We're Working On

  • Red-teaming & threat detection — Systematic methods to surface vulnerabilities and unsafe behaviors in deployed models.
  • Prompt-injection attacks & defenses — New injection techniques and defenses that separate trusted instructions from untrusted content.
  • Adversarial robustness — Training and evaluation methods that keep models reliable under adversarial conditions.

Related Publications

No publications in this area yet. Check back soon or view all research.