Robots Atlas>ROBOTS ATLAS
Architecture

Single-Stream DiT

2022ActivePublished: 1 October 2026Updated: 1 October 2026Published
Key innovation
Processing tokens of different modalities (text conditioning and image latents) within a single joint sequence handled by the same shared transformer blocks — instead of separate per-modality paths with dedicated weights, as in dual-stream MMDiT.
Category
Architecture
Abstraction level
Pattern
Operation level
ModelArchitecture block
Use cases
Text-to-image generationLatent diffusion modelsImage generation and editingDiffusion-based video generationConditional multimodal generation

How it works

1) Inputs from different modalities are tokenized: the image is encoded into latents (typically via a VAE) and split into patches, while text conditioning comes from a text encoder. 2) Text and image tokens are concatenated into one sequence. 3) The sequence passes through a stack of identical transformer blocks with shared weights; each block applies full self-attention over the entire joint sequence, so image and text tokens interact directly. 4) Global conditioning (diffusion timestep, optionally a class/pooled-text vector) is injected via adaLN modulation (scale/shift/gate) shared across all tokens. 5) The block outputs drive noise/velocity prediction in the diffusion denoising process, run iteratively over many steps.

Problem solved

Dual-stream architectures (MMDiT) maintain separate weights for each modality, which increases parameter count and complicates block design. Single-Stream DiT simplifies the architecture: a single set of weights processes all tokens, reducing per-block parameters, unifying information flow, and enabling full bidirectional interaction between tokens of all modalities within one self-attention operation.

Components

Joint token sequenceInput representation unifying modalities

Concatenation of text-conditioning tokens and image latent patches into one sequence that is the shared input to the whole block stack.

Shared transformer blocksCore computation

A stack of identical blocks (self-attention + MLP) with a single set of weights applied to all tokens regardless of modality — the key difference from dual-stream MMDiT.

Full self-attention over the joint sequenceCross-modal information mixing

A single attention operation spans text and image tokens at once, enabling bidirectional interaction without a separate cross-attention layer.

Official

adaLN modulation (conditioning)Global conditioning injection

Adaptive layer normalization producing scale, shift, and gate parameters from the conditioning vector (diffusion timestep, optionally pooled text/class), shared across the sequence.

Official

Implementation

Implementation pitfalls
Positional encoding in the joint sequenceMedium

Text and image with different structure share one sequence, so the positional-encoding scheme (e.g. RoPE, modality-specific offsets) must correctly distinguish and localize tokens.

Quadratic attention cost grows with sequence lengthHigh

Appending text tokens to the image sequence lengthens n and raises the O(n²) self-attention cost — significant for long prompts and high resolution.

Modality balance with shared weightsMedium

A single set of weights for both modalities may favor the dominant modality without proper conditioning and tuning, unlike the dedicated paths in MMDiT.

Evolution

Original paper · 2022 · ICCV 2023 · William Peebles
Scalable Diffusion Models with Transformers
William Peebles, Saining Xie
2022
Diffusion Transformer (DiT)
Inflection point

Peebles and Xie replace U-Net with a transformer operating on latent patches — a single sequence of image tokens conditioned via adaLN (the original, single-stream form of DiT).

2024
MMDiT (dual-stream) in Stable Diffusion 3
Inflection point

Introduces separate weights for the text and image modalities with joint attention — a dual-stream architecture against which single-stream is the simpler alternative.

2024
FLUX — hybrid double + single stream construction

FLUX combines double-stream blocks (DoubleStreamBlock) with single-stream blocks (SingleStreamBlock), in which all tokens flow through shared attention and MLP layers.

2025
Lumina-Image 2.0 — unified Next-DiT

Unified Next-DiT treats text and image tokens as a joint sequence, realizing fully single-stream multimodal processing.

Hyperparameters (configurable axes)

Number of blocks (depth)High

Number of shared transformer blocks in the stack.

Hidden sizeHigh

Token representation dimension; the primary model-scale driver.

Number of attention headsMedium

Number of heads in self-attention over the joint sequence.

Latent patch sizeMedium

Patch size when patchifying latents — affects the image token count, hence sequence length and attention cost.

Computational complexity

Time complexity: O(n² · d). Space complexity: O(n² + n · d).

Compute bottleneck

Full self-attention over the lengthened joint sequence

Because text and image share one sequence, length n is the sum of both modalities' tokens, increasing the quadratic attention cost relative to processing image patches alone.

Execution paradigm

Primary mode
Dense

Dense transformer: all weights active for all tokens, no per-modality routing.

Activation pattern
All paths active
Routing mechanism

Parallelism

Parallelism level
Partially parallel

Token processing within a block is fully parallel, but diffusion sampling at inference requires sequential denoising steps.

Scope
TrainingAcross tokens

Hardware requirements

Primary

Dense matrix multiplications in self-attention and MLP over the joint sequence efficiently utilize GPU tensor cores.