Robots Atlas>ROBOTS ATLAS
Gemini 1.5 Pro

Gemini 1.5 Pro

Gemini 1.5 Pro · Family: Gemini
Google DeepMind’s multimodal model (2024) with an MoE architecture and a 1–2M-token context window; natively handles text, image, audio and video. A legacy generation.
⚠ Deprecated⏳ Limited accessLLMMultimodalTool-using model📁 Gemini
Context window
1M (do 2M)
tokens
Release date
15 February 2024
Access:APIHostedDeployment:☁ Cloud

Overview

Gemini 1.5 Pro is a multimodal language model from Google DeepMind, introduced on 15 February 2024 as the flagship of the Gemini 1.5 generation. It stood out for its very long context window and native multimodality.

The model uses a Mixture-of-Experts (MoE) architecture that selectively activates the most relevant expert pathways depending on the input, improving efficiency. It supports a context window of up to 1 million tokens (up to 2 million in select variants; up to 10 million in testing) — allowing it to process, for example, an hour of video, 11 hours of audio, over 30,000 lines of code or over 700,000 words in a single pass.

Gemini 1.5 Pro is natively multimodal — it accepts text, images, audio and video (and code), and produces text responses (including code and structured data). It achieved results comparable to Gemini 1.0 Ultra while using less compute.

The model was available to developers and enterprises via the Gemini API, Google AI Studio and Google Vertex AI. Gemini 1.5 Pro is now a previous-generation (discontinued) model, superseded by newer Gemini models.

Classification
LLMMultimodalTool-using model
Family: Gemini
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 1M (do 2M)
Tools · ✓ Fine-tuning
📥 Input: text, image, audio, video
Platforms

Technical specification

Context window
1M (do 2M)
tokens
License
Proprietary
Features:Tool useFine-tuning
Modalities
⬇ Input
textimageaudiovideodocuments
⬆ Output
textcodestructured_data

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Audio understanding
Category: audio
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
OCR
Recognising text within images and documents.
Category: vision
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Multimodal understanding
Category: multimodal

Technical architecture

Deployment and security

☁ Available on platforms
🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 24 Jul 2026↗ Security documentation