Robots Atlas>ROBOTS ATLAS
Architecture

SwiGLU

2020ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Replacing the sigmoid gate of a Gated Linear Unit (GLU) with the Swish (SiLU) function in the Transformer feed-forward sublayer, improving model quality at a constant parameter budget.
Category
Architecture
Abstraction level
Primitive
Operation level
LayerArchitecture block
Use cases
Feed-forward sublayers in large language modelsTransformer architectures (encoder-decoder and decoder-only)Language model pre-training (T5, PaLM, LLaMA)Drop-in replacement for ReLU/GELU activations in the FFN

How it works

The input x is projected by two independent matrices: W (gate) and V (value). The gate output Swish₁(xW) is multiplied element-wise by xV, and the result is projected by an output matrix W₂: FFN_SwiGLU(x) = (Swish₁(xW) ⊗ xV)·W₂, where Swish₁(z) = z·σ(z) (β=1, i.e. SiLU). To keep the three matrices (W, V, W₂) from increasing the parameter count relative to a two-matrix FFN, the hidden dimension d_ff is reduced by about 2/3 (e.g. to 2/3·4d in LLaMA). The layer usually omits bias terms.

Problem solved

The conventional Transformer FFN sublayer with ReLU or GELU leaves room to improve model quality without increasing the parameter count. SwiGLU provides a more expressive, gated transformation that raises representation quality and benchmark scores while keeping the parameter budget constant.

Components

Gate projection (W)Gate

Linear projection of the input whose output is passed through Swish (β=1, SiLU) and acts as the gate.

Value projection (V)Value

The second linear projection of the input, multiplied element-wise by the gate output.

Element-wise gating (⊗)Gating mechanism

Hadamard product of the gate output Swish₁(xW) and the value projection xV.

Output projection (W2)Output projection

Linear projection mapping the hidden layer back to the model dimension.

Implementation

Implementation pitfalls
Not reducing the hidden dimensionHigh

Adding the third (gate) matrix without shrinking d_ff by about 2/3 increases FFN parameters by ~50% relative to the conventional variant.

Fix:Set d_ff to about 2/3 of the baseline (e.g. 2/3·4d) to preserve the parameter budget.
Wrong Swish β parameterLow

SwiGLU uses Swish with β=1 (SiLU); using a different β or making it trainable deviates from the paper's definition.

Fix:Use SiLU (Swish₁), i.e. z·σ(z).
Hardware-unfriendly dimensionMedium

The 2/3·4d factor can yield a dimension not aligned to the multiples used by tensor cores.

Fix:Round d_ff to a multiple (e.g. 128/256) for efficient GEMMs.

Evolution

Original paper · 2020 · arXiv preprint arXiv:2002.05202 · Noam Shazeer
GLU Variants Improve Transformer
Noam Shazeer
2016
Gated Linear Units (GLU) introduced

Dauphin et al. introduce GLU for language modeling with gated convolutional networks using a sigmoid gate.

2017
Swish / SiLU activation function

Introduction of the Swish activation (z·σ(βz)); its β=1 variant (SiLU) is used in the SwiGLU gate.

2020
SwiGLU introduced
Inflection point

Noam Shazeer introduces SwiGLU as a GLU variant in the Transformer FFN and shows quality gains on T5/GLUE/SuperGLUE.

2022
Adoption in PaLM

PaLM uses SwiGLU as the FFN activation in a 540B-parameter model.

2023
Adoption in LLaMA
Inflection point

LLaMA replaces ReLU with SwiGLU and uses a 2/3·4d dimension, popularizing SwiGLU across open LLMs.

Hyperparameters (configurable axes)

Hidden dimension (d_ff)High

FFN hidden dimension. Reduced by roughly 2/3 relative to a conventional FFN to offset the third matrix.

2048
2/3·4·d_model
Swish β parameterMedium

Parameter of the Swish function; set to 1 in SwiGLU (Swish₁ = SiLU).

1
Bias termsLow

Whether the projections use bias terms. In practice usually disabled (PaLM, LLaMA).

disabled

Computational complexity

Time complexity: O(n · d_model · d_ff). Space complexity: O(d_model · d_ff).

Compute bottleneck

Dense matrix multiplications (GEMM)

Performance dominated by three dense GEMM operations (W, V, W₂); no sparsity or routing.

Execution paradigm

Primary mode
Dense

All parameters are active for every token (no conditional computation).

Activation pattern
All paths active
Routing mechanism

Parallelism

Parallelism level
Fully parallel

The layer is fully parallel across tokens and within matmuls; no sequential dependencies.

Scope
TrainingInference

Hardware requirements

Primary

The three dense matrix multiplications map ideally onto GPU tensor cores.

Good fit

Dense GEMMs use the TPU matrix units well; the original T5 experiments ran on TPUs.