
Stable Diffusion 3.5
Stability AI text-to-image model (MMDiT): Large 8.1B, Large Turbo and Medium 2.5B variants; open weights under the Stability AI Community License.
Parameters
8.1B (Large) / 2.5B (Medium)
parameters
Release date
22 October 2024
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud
Overview
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open weights
Key parameters
🧩 Parameters: 8.1B (Large) / 2.5B (Medium)
✓ Fine-tuning
📥 Input: text
Technical specification
Parameters
8.1B (Large) / 2.5B (Medium)
parameters
License
Stability AI Community License
Hardware requirements
Open-weights models (Stability AI Community License) on Hugging Face; inference code on GitHub. The Medium variant runs on consumer GPUs (~9.9 GB VRAM excluding text encoders), while Large needs a larger-VRAM GPU. Also available via the Stability AI API and platforms (Replicate, Fireworks AI, DeepInfra, ComfyUI).
Features:✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
image
Capabilities and applications
Native model capabilities
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Text rendering in images
Generating images containing legible, correctly spelled text — infographics, posters, menu cards, QR codes, captions in a specific graphic style. A key capability that separates new-generation models from early image generators.
Category: vision
Reference-guided generation
Creating images based on previously supplied visual references — a specific person, artistic style, product, or space — preserving likeness and characteristic features.
Category: vision
Technical architecture
Core Architecture
Training Techniques