Robots Atlas>ROBOTS ATLAS
Architecture

Residual Connection

2015ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
A shortcut (skip) that adds a layer’s input to its output so the layer learns the residual — the foundation of training very deep networks.
Category
Architecture
Abstraction level
Building block
Operation level
Architecture blockLayer
Use cases
Deep convolutional networks (ResNet)Transformer blocks (attention + MLP)Training hundreds-of-layers networksStabilising gradient flow

How it works

Instead of learning the target mapping H(x), the block learns the residual F(x) = H(x) − x, and the output is F(x) + x. The identity shortcut provides a path with derivative 1, so gradients do not vanish when backpropagating through many layers. In Transformers the shortcut wraps sublayers (attention, MLP), usually with normalisation (pre-/post-LN).

Problem solved

Very deep networks without shortcuts suffer from vanishing gradients and degradation — harder to optimise than shallow ones. The residual connection unblocks gradient flow.

Components

Identity shortcutGradient highway

A path adding the input to the block output without transformation.

Residual branch F(x)Learned function

The transformation that learns only the residual relative to identity.

Evolution

Original paper · 2015 · Kaiming He
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun
2015
ResNet introduces residual connections and wins ImageNet
Inflection point
2017
Transformers adopt residual shortcuts around attention and MLP
2024
Hyper-Connections generalise residual connections