Robots Atlas>ROBOTS ATLAS
Multimodal

Text-to-Image

2021ActivePublished: 1 October 2026Updated: 1 October 2026Published
Key innovation
Reframes image creation as generation conditioned on a natural-language description — the model learns a shared text–image space, so any prompt becomes a steerable input to the generator.
Category
Multimodal
Abstraction level
Paradigm
Operation level
ModelInference
Use cases
Concept art and illustrationMarketing and visual designProduct prototyping and mockupsGame and film assetsImage editing and inpaintingSynthetic data for training vision models

How it works

1) A text encoder (CLIP/T5) turns the prompt into conditioning vectors. 2) A generative model produces the image: a diffusion approach starts from noise and iteratively denoises it over N steps, while a transformer/autoregressive approach generates image tokens. 3) Text conditioning is typically injected via cross-attention (or, in MMDiT, via joint processing of text and image tokens). 4) Classifier-Free Guidance (CFG) strengthens prompt adherence by interpolating conditional and unconditional predictions with a guidance scale. 5) In Latent Diffusion the whole process runs in an autoencoder's latent space, and a VAE decoder maps the latent back to the final full-resolution image.

Problem solved

Lets users create and edit visual content without drawing skills or production resources — a text description replaces manual illustration. It addresses controllable, open-ended image generation from an arbitrary prompt (zero-shot), rather than being limited to narrow domains or rigid templates.

Components

Text encoderSemantic conditioning

Converts the prompt into a representation that conditions the generator. Both contrastive (CLIP) and language (T5) encoders are used; Imagen showed the effectiveness of large text-only T5 encoders.

Official

Generative modelImage synthesis

The core that synthesizes the image. Today most often a diffusion backbone (U-Net or Diffusion Transformer / MMDiT); historically a GAN or an autoregressive transformer over image tokens.

Official

Conditioning (cross-attention)Prompt control

Binds the text representation to the generation process — usually via cross-attention between image tokens and text embeddings; in MMDiT, text and image are processed by a joint transformer with bidirectional information flow.

Official

Guidance (CFG)Text-adherence amplification

A technique that boosts image–prompt alignment by interpolating conditional and unconditional predictions, controlled by a guidance scale. Popularized for text-to-image in GLIDE.

Official

Implementation

Implementation pitfalls
Guidance scale too highMedium

Excessive CFG oversaturates colors, introduces artifacts, and reduces output diversity.

Fix:Tune the scale empirically per model; consider dynamic/rescaled CFG.
Too few sampling stepsMedium

Too few steps yield blurry or incoherent images with standard samplers.

Fix:Increase steps or use distilled samplers/models (few-step, rectified flow).
Poor prompt adherence (attribute binding)High

Models confuse attribute-to-object binding, object counts, and spatial relations.

Fix:Stronger text encoder (T5), better training captions, control techniques (regional/attention guidance).
Bias and content safetyHigh

Models inherit dataset biases and may generate unsafe or infringing content.

Fix:Data filtering, safety classifiers, watermarking, and usage policies.

Evolution

Original paper · 2021 · ICML 2021 · Aditya Ramesh
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, Ilya Sutskever
2018
AttnGAN — attention-based GAN text-to-image

Attentional GAN with multi-stage refinement driven by prompt words (arXiv 2017, CVPR 2018).

2021
DALL·E — autoregressive transformer over image tokens
Inflection point

Zero-shot text-to-image modeling a single stream of text and image tokens (dVAE + transformer).

2021
GLIDE — diffusion with Classifier-Free Guidance

Shows classifier-free guidance outperforming CLIP-based guidance for text-to-image.

2022
Latent Diffusion / Stable Diffusion
Inflection point

Moving diffusion into an autoencoder's latent space; the basis of Stable Diffusion (CVPR 2022).

2022
DALL·E 2 (unCLIP) and Imagen

DALL·E 2 generates images from CLIP latents; Imagen uses a large T5 text encoder. Both raise diffusion photorealism.

2022
Diffusion Transformer (DiT)
Inflection point

Replacing the U-Net backbone with a transformer over latent patches; strong scaling properties (FID vs GFLOPs).

2024
Stable Diffusion 3 (MMDiT) and FLUX.1

MMDiT with separate weights for text and image and bidirectional information flow; rectified flow. FLUX.1 by Black Forest Labs (August 2024).

2025
Qwen-Image — MMDiT foundation model

An MMDiT image-generation foundation model notable for complex text rendering and precise editing (technical report, August 2025).

Hyperparameters (configurable axes)

Guidance scale (CFG)Critical

Strength of prompt steering. Higher values increase text adherence but can reduce diversity and naturalness.

Number of sampling stepsHigh

Number of denoising iterations. More steps usually means better quality at the cost of time; the dominant factor in diffusion inference cost.

ResolutionHigh

Target image size (e.g. 512, 1024). Scales quadratically with pixels/tokens, strongly affecting memory and time.

Compute bottleneck

Iterative sampling loop

Repeated network passes in the denoising loop (N steps) are the main inference cost of diffusion models; reducing steps (distillation, rectified flow) is an active research direction.

Parallelism

Parallelism level
Partially parallel

Spatial computation within a single step is highly parallel (GPU), but successive denoising steps are sequential.

Scope
TrainingInference

Hardware requirements

Primary

Dense matrix ops in U-Net/DiT and the sampling loop benefit from tensor cores and high memory bandwidth.