Research Agenda
Shiba AI studies the safety of increasingly capable, interactive, and autonomous AI systems.
As AI evolves from isolated language models toward long-running interactions and multi-agent systems, important risks may emerge not from a single response, but over time, through interaction, and through the collective behavior of multiple AI agents.
Our research seeks to understand how these behaviors emerge, from the internal mechanisms of individual models to the dynamics of interacting agents, and to develop methods for detecting, controlling, and governing them.
Connecting the Agenda
These research directions are deeply connected. Research on AI behavior and multi-agent systems reveals new forms of emergent behavior. Mechanistic interpretability helps uncover the mechanisms behind these behaviors. Detection, control, and governance research then translates this understanding into practical approaches for monitoring and improving AI systems.
Together, these directions allow us to study AI safety from internal model mechanisms to interaction-level dynamics and, ultimately, to system-level intervention and governance.
Long-Term Goal
Our long-term objective is to move AI safety research beyond simply observing failures toward a deeper understanding of how and why they occur:
Through this research agenda, Shiba AI aims to contribute both fundamental scientific understanding and practical approaches for building trustworthy AI systems.
Research Areas
Exploring critical sub-sections of AI alignment and safety research.