About
Moshi is a family of real-time voice AI models developed by Kyutai (a Paris AI lab). It includes the full-duplex speech-to-speech Moshi model and its synthetic-data fine-tuned variants: Moshiko (male voice) and Moshika (female voice). The models combine the Helium (7B) language backbone with the Mimi neural audio codec and the “Inner Monologue” method.

