Robots Atlas>ROBOTS ATLAS
FLUX.1

FLUX.1

FLUX.1 ([pro] / [dev] / [schnell], 12B) · Family: FLUX
Black Forest Labs text-to-image model (12B, rectified-flow transformer): [pro]/[dev]/[schnell] variants; quality comparable to DALL·E 3 and Midjourney 6.
✓ Active✓ Public access⚖ Open weightsImage generation📁 FLUX
Parameters
12B
parameters
Release date
1 August 2024
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud

Overview

FLUX.1 is an image generation model developed by Black Forest Labs and released in August 2024. Black Forest Labs is a German startup from Freiburg, founded by former Stability AI employees: Robin Rombach, Andreas Blattmann and Patrick Esser.

The model generates images from natural-language prompts (text-to-image) and also supports image-to-image transformations. It is based on rectified-flow transformer blocks (trained with flow matching) scaled to 12 billion parameters — a shift from classic diffusion toward flow matching.

FLUX.1 came in three variants differing in license and access: [schnell] (open source, Apache 2.0, speed-optimized), [dev] (source-available, non-commercial license, with commercial licensing available) and [pro] (proprietary, available via API only and licensed to third parties).

In testing, FLUX.1 matches leading models: in prompt fidelity it is comparable to DALL·E 3, and in photorealism to Midjourney 6; it also handles hand generation better than earlier models (e.g. Stable Diffusion XL). Regardless of the chosen variant, users retain ownership of the generated images.

Classification
Image generation
Family: FLUX
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open weights
Key parameters
🧩 Parameters: 12B
✓ Fine-tuning
📥 Input: text, image

Technical specification

Parameters
12B
parameters
License
[schnell]: Apache 2.0; [dev]: licencja niekomercyjna (source-available); [pro]: własnościowa (API)
Hardware requirements
The [schnell] and [dev] variants have open/released weights (Hugging Face) and can be run locally on a GPU with sufficient VRAM (12B parameters; quantized versions reduce requirements). The [pro] variant is available via API only (e.g. Black Forest Labs / partners).
Features:Fine-tuning
Modalities
⬇ Input
textimage
⬆ Output
image

Capabilities and applications

Native model capabilities
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision
Text rendering in images
Generating images containing legible, correctly spelled text — infographics, posters, menu cards, QR codes, captions in a specific graphic style. A key capability that separates new-generation models from early image generators.
Category: vision
Reference-guided generation
Creating images based on previously supplied visual references — a specific person, artistic style, product, or space — preserving likeness and characteristic features.
Category: vision

Technical architecture