Robots Atlas>ROBOTS ATLAS
Stable Diffusion 3.5

Stable Diffusion 3.5

Stable Diffusion 3.5 (Large 8.1B / Large Turbo / Medium 2.5B) · Family: Stable Diffusion
Stability AI text-to-image model (MMDiT): Large 8.1B, Large Turbo and Medium 2.5B variants; open weights under the Stability AI Community License.
✓ Active✓ Public access⚖ Open weightsImage generation📁 Stable Diffusion
Parameters
8.1B (Large) / 2.5B (Medium)
parameters
Release date
22 October 2024
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud

Overview

Stable Diffusion 3.5 is the latest generation of Stability AI’s text-to-image model family, announced on 22 October 2024 (the Medium variant was added on 29 October 2024). The models are based on an improved Multimodal Diffusion Transformer (MMDiT) architecture, including Query-Key normalization in the transformer blocks.

The family includes three main variants: Large (8.1 billion parameters, for professional use, ~1 megapixel resolution), Large Turbo (a distilled version generating high-quality images in 4 steps with excellent prompt adherence) and Medium (2.5 billion parameters, 0.25–2 megapixel resolution, optimized for consumer hardware — the Medium variant uses the MMDiT-X architecture).

The models stand out for strong prompt adherence, diversity of generated people (including skin tones and features) and stylistic versatility — from photography and painting to 3D. The Medium variant can run on consumer GPUs, requiring about 9.9 GB of VRAM (excluding text encoders).

Stable Diffusion 3.5 was released under the Stability AI Community License, which permits free non-commercial use and commercial use for organizations with under $1M in annual revenue; users retain ownership of generated images. Weights are available on Hugging Face, inference code on GitHub, and the model is also accessible via the Stability AI API and platforms such as Replicate, Fireworks AI, DeepInfra and ComfyUI.

Classification
Image generation
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open weights
Key parameters
🧩 Parameters: 8.1B (Large) / 2.5B (Medium)
✓ Fine-tuning
📥 Input: text

Technical specification

Parameters
8.1B (Large) / 2.5B (Medium)
parameters
License
Stability AI Community License
Hardware requirements
Open-weights models (Stability AI Community License) on Hugging Face; inference code on GitHub. The Medium variant runs on consumer GPUs (~9.9 GB VRAM excluding text encoders), while Large needs a larger-VRAM GPU. Also available via the Stability AI API and platforms (Replicate, Fireworks AI, DeepInfra, ComfyUI).
Features:Fine-tuning
Modalities
⬇ Input
text
⬆ Output
image

Capabilities and applications

Native model capabilities
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Text rendering in images
Generating images containing legible, correctly spelled text — infographics, posters, menu cards, QR codes, captions in a specific graphic style. A key capability that separates new-generation models from early image generators.
Category: vision
Reference-guided generation
Creating images based on previously supplied visual references — a specific person, artistic style, product, or space — preserving likeness and characteristic features.
Category: vision

Technical architecture