Single-Stream DiT
How it works
1) Inputs from different modalities are tokenized: the image is encoded into latents (typically via a VAE) and split into patches, while text conditioning comes from a text encoder. 2) Text and image tokens are concatenated into one sequence. 3) The sequence passes through a stack of identical transformer blocks with shared weights; each block applies full self-attention over the entire joint sequence, so image and text tokens interact directly. 4) Global conditioning (diffusion timestep, optionally a class/pooled-text vector) is injected via adaLN modulation (scale/shift/gate) shared across all tokens. 5) The block outputs drive noise/velocity prediction in the diffusion denoising process, run iteratively over many steps.
Problem solved
Dual-stream architectures (MMDiT) maintain separate weights for each modality, which increases parameter count and complicates block design. Single-Stream DiT simplifies the architecture: a single set of weights processes all tokens, reducing per-block parameters, unifying information flow, and enabling full bidirectional interaction between tokens of all modalities within one self-attention operation.
Components
Concatenation of text-conditioning tokens and image latent patches into one sequence that is the shared input to the whole block stack.
A stack of identical blocks (self-attention + MLP) with a single set of weights applied to all tokens regardless of modality — the key difference from dual-stream MMDiT.
A single attention operation spans text and image tokens at once, enabling bidirectional interaction without a separate cross-attention layer.
Official
Adaptive layer normalization producing scale, shift, and gate parameters from the conditioning vector (diffusion timestep, optionally pooled text/class), shared across the sequence.
Official
Implementation
Text and image with different structure share one sequence, so the positional-encoding scheme (e.g. RoPE, modality-specific offsets) must correctly distinguish and localize tokens.
Appending text tokens to the image sequence lengthens n and raises the O(n²) self-attention cost — significant for long prompts and high resolution.
A single set of weights for both modalities may favor the dominant modality without proper conditioning and tuning, unlike the dedicated paths in MMDiT.
Evolution
Peebles and Xie replace U-Net with a transformer operating on latent patches — a single sequence of image tokens conditioned via adaLN (the original, single-stream form of DiT).
Introduces separate weights for the text and image modalities with joint attention — a dual-stream architecture against which single-stream is the simpler alternative.
FLUX combines double-stream blocks (DoubleStreamBlock) with single-stream blocks (SingleStreamBlock), in which all tokens flow through shared attention and MLP layers.
Unified Next-DiT treats text and image tokens as a joint sequence, realizing fully single-stream multimodal processing.
Hyperparameters (configurable axes)
Number of shared transformer blocks in the stack.
Token representation dimension; the primary model-scale driver.
Number of heads in self-attention over the joint sequence.
Patch size when patchifying latents — affects the image token count, hence sequence length and attention cost.
Computational complexity
Time complexity: O(n² · d). Space complexity: O(n² + n · d).
Compute bottleneck
Because text and image share one sequence, length n is the sum of both modalities' tokens, increasing the quadratic attention cost relative to processing image patches alone.
Execution paradigm
Dense transformer: all weights active for all tokens, no per-modality routing.
Parallelism
Token processing within a block is fully parallel, but diffusion sampling at inference requires sequential denoising steps.
Hardware requirements
Dense matrix multiplications in self-attention and MLP over the joint sequence efficiently utilize GPU tensor cores.