Robots Atlas>ROBOTS ATLAS
Gemini 3.8 Flash

Gemini 3.8 Flash

3.8 Flash · Family: Gemini
Gemini Flash-family model (Google DeepMind) for agentic tasks, software engineering and complex workflows; 1M-token context window.
✓ Active✓ Public accessLLMMultimodalReasoning modelTool-using model📁 Gemini
Context window
1M
tokens
Max output
64,000
tokens
Access:APIHostedDeployment:☁ Cloud

Overview

Gemini 3.8 Flash is a model in the Gemini Flash family developed by Google DeepMind, positioned as a workhorse model for agentic tasks, software engineering and complex enterprise workflows. It provides advanced reasoning at the latency and scale typical of Flash models. It accepts text, image, video, audio and PDF input, with a 1M-token context window and a maximum output of 64k tokens. It supports function calling, search as a tool and computer use. The model is generally available via the API and Google hosted interfaces.

Classification
LLMMultimodalReasoning modelTool-using model
Family: Gemini
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 1M
✓ Tools
📥 Input: text, image, video, audio…

Technical specification

Context window
1M
tokens
Max output tokens
64,000
tokens per response
Features:✓ Tool use
Modalities
⬇ Input
textimagevideoaudiodocuments
⬆ Output
text

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Multimodal understanding
Category: multimodal
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Audio understanding
Category: audio
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video Understanding
Category: video
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
OCR
Recognising text within images and documents.
Category: vision
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Planning
Forming and executing action plans for complex tasks.
Category: planning
Interleaved Multimodal Input
Category: reasoning
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Computer use
The model's ability to operate a computer interface by interpreting screenshots and generating actions such as clicks, typing, and navigating applications.
Category: planning

Benchmark results

4 benchmarks
DeepSWE v1.1
success rate
>70%
📄 website
High success at low cost (official DeepMind data).
Vals Finance Agent v2
61.4%%
📄 website
Harvey's Legal Agent Benchmark
10.0%%
📄 website
HLE-Verified
accuracy
54.9%%
📄 website

Pricing

Technical architecture