Google DeepMind's video generation model (2025) with native audio (dialogue, effects, ambience). 8-second clips, up to 4K resolution, realistic physics; Standard and Fast variants; SynthID watermark.
Release date
20 May 2025
Access:APIHostedDeployment:โ Cloud
Overview
Applications
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
๐ฅ Input: text, image
Technical specification
License
Wlasnosciowa (zamkniete wagi)
Hardware requirements
Cloud-only access (Gemini API / Vertex AI); no downloadable weights. 8-second clips, up to 4K.
Modalities
โฌ Input
textimage
โฌ Output
videoaudio
Capabilities and applications
Native model capabilities
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Text-to-video generation
Generating video sequences from a text prompt (and optionally an image), with coherent scene dynamics.
Category: video
Image-to-video
The model's ability to animate a static input image โ extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Multi-shot video generation
The model's ability to generate narratively coherent sequences composed of multiple shots, preserving character identity, style and continuity across shot transitions.
Category: video
Native audio generation for video
A video model's ability to natively generate an audio track together with the visuals: dialogue, sound effects and ambient noise synchronized with the generated video.
Category: multimodal
Pricing
Technical architecture
Core Architecture
Deployment and security
๐ Security / Enterprise
โ Verified enterprise information
Updated: 30 Jul 2026โ Security documentation
