Text-to-Image
How it works
1) A text encoder (CLIP/T5) turns the prompt into conditioning vectors. 2) A generative model produces the image: a diffusion approach starts from noise and iteratively denoises it over N steps, while a transformer/autoregressive approach generates image tokens. 3) Text conditioning is typically injected via cross-attention (or, in MMDiT, via joint processing of text and image tokens). 4) Classifier-Free Guidance (CFG) strengthens prompt adherence by interpolating conditional and unconditional predictions with a guidance scale. 5) In Latent Diffusion the whole process runs in an autoencoder's latent space, and a VAE decoder maps the latent back to the final full-resolution image.
Problem solved
Lets users create and edit visual content without drawing skills or production resources — a text description replaces manual illustration. It addresses controllable, open-ended image generation from an arbitrary prompt (zero-shot), rather than being limited to narrow domains or rigid templates.
Components
Converts the prompt into a representation that conditions the generator. Both contrastive (CLIP) and language (T5) encoders are used; Imagen showed the effectiveness of large text-only T5 encoders.
Official
The core that synthesizes the image. Today most often a diffusion backbone (U-Net or Diffusion Transformer / MMDiT); historically a GAN or an autoregressive transformer over image tokens.
Official
Binds the text representation to the generation process — usually via cross-attention between image tokens and text embeddings; in MMDiT, text and image are processed by a joint transformer with bidirectional information flow.
Official
A technique that boosts image–prompt alignment by interpolating conditional and unconditional predictions, controlled by a guidance scale. Popularized for text-to-image in GLIDE.
Official
Implementation
Excessive CFG oversaturates colors, introduces artifacts, and reduces output diversity.
Too few steps yield blurry or incoherent images with standard samplers.
Models confuse attribute-to-object binding, object counts, and spatial relations.
Models inherit dataset biases and may generate unsafe or infringing content.
Evolution
Attentional GAN with multi-stage refinement driven by prompt words (arXiv 2017, CVPR 2018).
Zero-shot text-to-image modeling a single stream of text and image tokens (dVAE + transformer).
Shows classifier-free guidance outperforming CLIP-based guidance for text-to-image.
Moving diffusion into an autoencoder's latent space; the basis of Stable Diffusion (CVPR 2022).
DALL·E 2 generates images from CLIP latents; Imagen uses a large T5 text encoder. Both raise diffusion photorealism.
Replacing the U-Net backbone with a transformer over latent patches; strong scaling properties (FID vs GFLOPs).
MMDiT with separate weights for text and image and bidirectional information flow; rectified flow. FLUX.1 by Black Forest Labs (August 2024).
An MMDiT image-generation foundation model notable for complex text rendering and precise editing (technical report, August 2025).
Hyperparameters (configurable axes)
Strength of prompt steering. Higher values increase text adherence but can reduce diversity and naturalness.
Number of denoising iterations. More steps usually means better quality at the cost of time; the dominant factor in diffusion inference cost.
Target image size (e.g. 512, 1024). Scales quadratically with pixels/tokens, strongly affecting memory and time.
Compute bottleneck
Repeated network passes in the denoising loop (N steps) are the main inference cost of diffusion models; reducing steps (distillation, rectified flow) is an active research direction.
Parallelism
Spatial computation within a single step is highly parallel (GPU), but successive denoising steps are sequential.
Hardware requirements
Dense matrix ops in U-Net/DiT and the sampling loop benefit from tensor cores and high memory bandwidth.