Robots Atlas>ROBOTS ATLAS
Gemini 3.6 Flash

Gemini 3.6 Flash

3.6 Flashย ยทย Family: Gemini
A previous-generation model in the Gemini Flash line (Google DeepMind), balancing speed and multimodal capabilities across general agentic and everyday tasks.
โœ“ Activeโœ“ Public accessLLMMultimodalReasoning modelTool-using model๐Ÿ“ Gemini
Context window
1M
tokens
Max output
65,536
tokens
Access:APIHostedDeployment:โ˜ Cloud

Overview

Gemini 3.6 Flash is a previous-generation model in the Gemini Flash line, developed by Google DeepMind. Marked as "Stable", it balances speed and multimodal capabilities across general agentic and everyday tasks, while retaining the low latency and high throughput characteristic of Flash variants.

The model accepts text, images, audio, video and PDF documents as input and produces text and code as output. It offers a 1,048,576-token input context window and up to 65,536 output tokens. It supports thinking, function calling, structured output, code execution, computer use (in Preview), grounding with Google Search and Google Maps, Batch API processing and context caching. The Live API and image/audio generation are not supported. Its API model ID is gemini-3.6-flash.

It is available through the Gemini API, Google AI Studio, Vertex AI, Google Antigravity and Gemini Enterprise, among others. Latest update: July 2026. The knowledge cutoff date is not stated in the official documentation.

Classification
LLMMultimodalReasoning modelTool-using model
Family: Gemini
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
๐Ÿ“ Context: 1M
โœ“ Tools
๐Ÿ“ฅ Input: text, image, audio, videoโ€ฆ

Technical specification

Context window
1M
tokens
Max output tokens
65,536
tokens per response
License
proprietary
Hardware requirements
Available only through Google cloud infrastructure (Gemini API, Vertex AI, Google AI Studio).
Features:โœ“ Tool use
Modalities
โฌ‡ Input
textimageaudiovideodocuments
โฌ† Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Long context
Support for large context windows โ€” tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Multimodal understanding
Category: multimodal
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Audio understanding
Category: audio
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video Understanding
Category: video
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
OCR
Recognising text within images and documents.
Category: vision
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Planning
Forming and executing action plans for complex tasks.
Category: planning
Interleaved Multimodal Input
Category: reasoning

Technical architecture