Black Forest Labs text-to-image model (12B, rectified-flow transformer): [pro]/[dev]/[schnell] variants; quality comparable to DALL·E 3 and Midjourney 6.
Parameters
12B
parameters
Release date
1 August 2024
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud
Overview
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open weights
Key parameters
🧩 Parameters: 12B
✓ Fine-tuning
📥 Input: text, image
Technical specification
Parameters
12B
parameters
License
[schnell]: Apache 2.0; [dev]: licencja niekomercyjna (source-available); [pro]: własnościowa (API)
Hardware requirements
The [schnell] and [dev] variants have open/released weights (Hugging Face) and can be run locally on a GPU with sufficient VRAM (12B parameters; quantized versions reduce requirements). The [pro] variant is available via API only (e.g. Black Forest Labs / partners).
Features:✓ Fine-tuning
Modalities
⬇ Input
textimage
⬆ Output
image
Capabilities and applications
Native model capabilities
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision
Text rendering in images
Generating images containing legible, correctly spelled text — infographics, posters, menu cards, QR codes, captions in a specific graphic style. A key capability that separates new-generation models from early image generators.
Category: vision
Reference-guided generation
Creating images based on previously supplied visual references — a specific person, artistic style, product, or space — preserving likeness and characteristic features.
Category: vision
Technical architecture
Core Architecture
Training Techniques
