Research Agenda

Shiba AI studies the safety of increasingly capable, interactive, and autonomous AI systems.

As AI evolves from isolated language models toward long-running interactions and multi-agent systems, important risks may emerge not from a single response, but over time, through interaction, and through the collective behavior of multiple AI agents.

Our research seeks to understand how these behaviors emerge, from the internal mechanisms of individual models to the dynamics of interacting agents, and to develop methods for detecting, controlling, and governing them.

AI Behavior, Interaction & Emergence

We study how safety-relevant behavior develops in both individual AI models and multi-agent systems.

  • At the individual level, our interests include reasoning behavior, bias, persona, hallucination, and context-dependent decision making.
  • At the multi-agent level, we investigate coordination, communication, network topology, consensus, polarization, collusion, and other forms of emergent collective behavior.

Central question — How can individually capable or apparently safe AI agents produce unexpected or unsafe behavior when they interact?

We aim to understand the conditions under which these behaviors emerge and how interaction structures can be designed to make multi-agent AI systems more reliable and safe.

Connecting the Agenda

These research directions are deeply connected. Research on AI behavior and multi-agent systems reveals new forms of emergent behavior. Mechanistic interpretability helps uncover the mechanisms behind these behaviors. Detection, control, and governance research then translates this understanding into practical approaches for monitoring and improving AI systems.

Together, these directions allow us to study AI safety from internal model mechanisms to interaction-level dynamics and, ultimately, to system-level intervention and governance.

Long-Term Goal

Our long-term objective is to move AI safety research beyond simply observing failures toward a deeper understanding of how and why they occur:

Understanding mechanisms
Predicting emergence
Detecting risks
Intervening
Governing

Through this research agenda, Shiba AI aims to contribute both fundamental scientific understanding and practical approaches for building trustworthy AI systems.

Research Areas

Exploring critical sub-sections of AI alignment and safety research.