Robots Atlas>ROBOTS ATLAS
Architecture

ViT

2020ActivePublished: 28 May 2026Updated: 28 May 2026Published
Key innovation
Applying a pure Transformer architecture to images by splitting them into a sequence of flat patches (e.g. 16×16 px) treated as tokens — showing that the inductive biases of convolutions (locality, translation equivariance) are not required if the model is pretrained on sufficiently large datasets.
Category
Architecture
Abstraction level
System
Operation level
ModelLayerArchitecture block
Use cases
Image classification (ImageNet, fine-grained)Vision backbone for CLIP, ALIGN, SigLIP (zero-shot)Self-supervised pretraining (DINO, DINOv2, MAE, iBOT, BEiT)Segmentation (SAM, Segment Anything Model)Object detection (DETR-style, ViTDet, OWL-ViT)Vision-Language Models (LLaVA, Flamingo, PaLI, Gemini Vision)Vision-Language-Action in robotics (RT-2, OpenVLA, π0)Medical imaging, satellite imagery, video understanding

How it works

Step 1 — Patching: an H × W × C image is split into N patches of P × P (e.g. 224 × 224 → 14 × 14 = 196 patches of 16 × 16). Implementation-wise this is a Conv2d(in=C, out=d_model, kernel=P, stride=P) — efficient on GPU. Step 2 — Patch embedding: each patch (P²·C dims) is linearly projected to d_model. Step 3 — [CLS] token: a learned d_model vector is prepended as the classification token (analogous to BERT). Step 4 — Positional embedding: a learned 1D position vector (length N+1) is added to each token to carry spatial information (self-attention itself is permutation-invariant). Step 5 — Transformer encoder stack: L layers, each with LayerNorm → Multi-Head Self-Attention → residual → LayerNorm → FFN (gelu) → residual. ViT uses pre-norm (LN before attention). Step 6 — Classification: the last [CLS] representation goes through an MLP / linear head → softmax over classes. Standard training uses supervised cross-entropy; modern variants use masked image modeling (MAE), contrastive learning (CLIP/DINO) or self-distillation (DINOv2) as pretraining. Inference at a new resolution requires interpolating positional embeddings.

Problem solved

How to reach state-of-the-art image classification without relying on the hand-engineered inductive biases of convolutions (locality, translation equivariance, hierarchical receptive fields), and how to unify NLP and vision architectures, enabling multimodal models with a single backbone.

Components

Patch embeddingImage tokenization

Split of the image into N non-overlapping P × P patches and their linear projection to d_model. Typically implemented as Conv2d(C, d_model, kernel=P, stride=P).

INImage tensor — batch, channels (typically 3 for RGB), height, width.
OUTSequence of N patch embeddings of dimension d_model.

Official

[CLS] tokenClassification placeholder aggregating information

A learned d_model vector prepended to the sequence. Its last-layer representation is used as the global image descriptor for classification.

Official

Positional embedding (1D learned)Injecting spatial information

A learned [N+1, d_model] tensor added to the tokens, since self-attention is permutation-invariant and does not know patch positions on its own.

1D learnedDefault in original ViT.
2D learnedSeparate embeddings for x and y axes.
SinusoidalStatic, as in NLP Transformer.
Relative / RoPEIntroduced in newer variants (e.g. ViT-22B).

Official

Transformer encoder block (pre-norm)Modeling global dependencies between patches

L layers — each: LN → MHSA → residual → LN → FFN(GELU) → residual. Identical to BERT/GPT, without a causal mask (all-to-all attention).

Classification headOutput

MLP or linear layer mapping the [CLS] representation to class logits. Replaced by a projection head in self-supervised pretraining (e.g. DINO MLP).

Official

Implementation

Implementation pitfalls
Poor results without large-scale pretrainingHigh

Training ViT from scratch on ImageNet-1k yields lower accuracy than ResNet — without the convolutional inductive biases the model needs far more data.

Fix:Pretrain on ImageNet-21k / JFT or distill (DeiT). Strong augmentation (RandAugment, Mixup, CutMix), stochastic depth.
Positional-embedding interpolation when changing resolutionHigh

Learned 1D positional embeddings are specific to the pretraining N. Fine-tuning at 384×384 after pretraining at 224×224 requires 2D interpolation, otherwise performance drops.

Fix:Reshape to 2D, bilinear / biquadratic interpolation, then flatten back. Or use RoPE / relative position.
Quadratic cost at high resolutionsHigh

Dense tasks (segmentation, detection) require high resolutions; standard ViT scales O(N²) in number of patches.

Fix:Swin (local windows), FlashAttention (better constant), adaptive tokens (Token Merging), hierarchical backbones.
No hierarchical receptive fieldsMedium

CNNs naturally build a hierarchy of feature maps from local to global; standard ViT operates at a single scale, which can be problematic for detection of differently sized objects.

Fix:Swin Transformer, MViT, PVT introduce a hierarchy. ViTDet shows that for detection a "simple feature pyramid" suffices.
Unstable training of large ViTsMedium

Very deep / large ViTs (ViT-H/22B) suffer from attention divergence in deep layers.

Fix:QK-norm (query/key normalization), more careful LN, gradient clipping, learning-rate warm-up, freezing patch embedding initially.

Evolution

Original paper · 2020 · ICLR 2021 · Alexey Dosovitskiy
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby
2017
Transformer (Vaswani et al.) — source architecture

Self-attention without recurrence emerges in NLP — the foundation of the later ViT.

2020
iGPT (Chen et al., OpenAI) — autoregressive generative pretraining on pixels

First influential demonstration of a pure Transformer on images (at the pixel level), a precursor to ViT.

2020
ViT — "An Image is Worth 16x16 Words" (Dosovitskiy et al.)
Inflection point

Full formulation of ViT: 16×16 patching, pure Transformer, pretraining on JFT-300M. ImageNet result beats the best CNNs.

2021
DeiT (Touvron et al., Meta) — data-efficient ViT
Inflection point

Shows that ViT can be trained on ImageNet-1k without massive pretraining via distillation and improved augmentation.

2021
Swin Transformer (Liu et al., Microsoft) — hierarchical windowed ViT
Inflection point

Local self-attention windows + shifted windows + multi-resolution hierarchy — make ViT competitive as a general-purpose backbone (detection, segmentation).

2021
CLIP (Radford et al., OpenAI) — ViT as the visual encoder in multimodality
Inflection point

ViT becomes the standard backbone for contrastive image-text pretraining; opens the era of zero-shot vision.

2021
MAE (He et al., Meta) — masked autoencoder pretraining for ViT
Inflection point

Masking ~75% of patches and reconstructing — a highly efficient self-supervised pretraining for ViT.

2021
DINO (Caron et al., Meta) — self-distillation with no labels

Self-supervised pretraining of ViT reveals emergent segmentation properties in attention maps.

2023
ViT-22B (Dehghani et al., Google) — scaling ViT to 22B parameters

Shows that ViT scales analogously to LLMs; reveals new behavioral properties at large scale.

2023
DINOv2 (Oquab et al., Meta) and SAM (Kirillov et al., Meta) — ViT as a universal backbone

ViT becomes the foundation of open vision foundation models: general-purpose features (DINOv2) and promptable segmentation (SAM).

Hyperparameters (configurable axes)

Patch size (P)Critical

Pixels per patch (16, 14, 8). Smaller P → more tokens → quadratically more expensive attention, but better spatial resolution.

16Default in ViT-B/16, ViT-L/16.
14DINOv2, used for dense segmentation.
8ViT-B/8 — very dense sequence, expensive.
Model sizeCritical

Standard variants: ViT-Ti, ViT-S, ViT-B (Base, ~86M), ViT-L (Large, ~307M), ViT-H (Huge, ~632M), ViT-g/G, ViT-22B.

Input resolutionHigh

Most commonly 224×224 (pretraining), 384×384 (fine-tuning). Changing resolution requires positional-embedding interpolation.

Number of attention headsHigh

Standard: 12 (ViT-B), 16 (ViT-L), 16 (ViT-H). Head dimension is d_model / num_heads.

Pretraining data scaleCritical

A critical axis from the original paper: ViT loses to ResNet on ImageNet-1k but wins on ImageNet-21k and JFT-300M.

Positional encoding typeMedium

1D learned (original), 2D learned, sinusoidal, relative, RoPE — affects the ability to change resolution.

Computational complexity

Time complexity: O(N² · d) + O(N · d²) per layer. Space complexity: O(N² + N · d).

Execution paradigm

Primary mode
Dense

Standard ViT is a dense model — all parameters active for every patch. MoE variants (V-MoE) introduce conditional computation but are not part of the core definition.

Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel

ViT is an encoder (no causal mask) — all patches are processed in parallel during both training and inference. Ideal for tensor and sequence parallelism in very large models (ViT-22B).

Scope
TrainingInferenceAcross tokensAcross devices

Hardware requirements

Primary

ViT is a dense Transformer — all ops (patch embedding, MHSA, FFN) map to matmul and are ideal for tensor cores (FP16/BF16/FP8).

Primary

ViT originated at Google on TPU and is trained there up to 22B scale; the systolic array handles MHSA and FFN excellently.

Possible

ViT-B/S inference on CPU AVX/AVX-512 (ONNX Runtime, OpenVINO) is practical for batch use cases, though slower than on GPU.

Limited

Academic FPGA accelerators for ViT exist, but a broad production ecosystem is lacking.