Kyutai’s open real-time voice model (full-duplex speech-to-speech): Helium 7B backbone + Mimi codec, ~160 ms latency.
Parameters
≈8B (backbone Helium 7B)
parameters
Release date
18 September 2024
Access:DownloadHostedDeployment:💻 Local📱 On-device☁ Cloud
Overview
Applications
Access & deployment
DownloadHosted
LocalOn-deviceCloud
Weights: Open source
Key parameters
🧩 Parameters: ≈8B (backbone Helium 7B)
✓ Fine-tuning
📥 Input: audio
Platforms
Technical specification
Parameters
≈8B (backbone Helium 7B)
parameters
License
Wagi CC-BY-4.0; kod Apache-2.0
Hardware requirements
Open source; runs on GPU, with quantized variants (MLX) also available to run locally/on-device (e.g. Apple Silicon).
Features:✓ Fine-tuning
Modalities
⬇ Input
audio
⬆ Output
audiotext
Capabilities and applications
Native model capabilities
Voice Conversation
Ability to conduct multi-turn real-time voice conversations with context retention and natural speech pacing.
Category: speech
Natural conversation
Conducting a conversation with a tone close to human: a warmer voice, empathy in emotional responses, humour, and avoiding the stiff 'AI assistant jargon'. Introduced as a deliberate improvement in GPT-5.1 Instant.
Category: language
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Speech to text
Category: speech
Text to speech
Category: speech
Streaming Speech-to-Text
Real-time conversion of speech to text with immediate output as the speaker is talking.
Category: speech
Application domains
Technical architecture
Core Architecture
Model Form
Training Techniques
Deployment and security
☁ Available on platforms
