Robots Atlas>ROBOTS ATLAS
Gemini 2.0 Flash

Gemini 2.0 Flash

Gemini 2.0 Flash · Family: Gemini
Google DeepMind’s fast workhorse model (2024/25): multimodal, 1M context, native image and speech (TTS) generation, native tools and Live API. A legacy generation.
⚠ Deprecated⏳ Limited accessLLMMultimodalTool-using modelImage generation📁 Gemini
Context window
1M
tokens
Release date
11 December 2024
Access:APIHostedDeployment:☁ Cloud

Overview

Gemini 2.0 Flash is a multimodal language model from Google DeepMind, introduced on 11 December 2024 (experimentally) and generally available from 30 January 2025. It was designed as an efficient, low-latency “workhorse” model, running about twice as fast as Gemini 1.5 Pro and outperforming it on key benchmarks.

The model accepts multimodal input (text, image, audio, video) and — distinctively — natively generates images and controllable, multilingual speech (TTS) with watermarking (SynthID). It supports a context window of up to 1 million tokens.

Gemini 2.0 Flash natively uses tools: Google Search, code execution and user-defined functions, and via the Multimodal Live API enables real-time audio/video interactions. It supports a new class of agentic experiences, combining multimodal reasoning, long context and precise instruction following.

The model was available to developers and enterprises via the Gemini API, Google AI Studio and Google Vertex AI. Gemini 2.0 Flash is now a previous-generation (discontinued) model, superseded by newer Gemini models (2.5 and 3).

Classification
LLMMultimodalTool-using modelImage generation
Family: Gemini
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 1M
Tools · ✓ Fine-tuning
📥 Input: text, image, audio, video
Platforms

Technical specification

Context window
1M
tokens
License
Proprietary
Features:Tool useFine-tuning
Modalities
⬇ Input
textimageaudiovideodocuments
⬆ Output
textcodeimageaudiostructured_data

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Image understanding
Analysing and interpreting the content of images.
Category: vision
Audio understanding
Category: audio
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Multimodal understanding
Category: multimodal
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Text to speech
Category: speech
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language

Technical architecture

Deployment and security

☁ Available on platforms
🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 24 Jul 2026↗ Security documentation