Architecture
Luong Attention
2015HistoricalPublished: 28 May 2026Updated: 28 May 2026Published
Key
innovation
Simplifying and systematising NMT attention through global and local variants and multiplicative/dot-product scoring instead of additive scoring.
Category
Architecture
Abstraction level
Building block
Operation level
Architecture blockTrainingInference
Use cases
Neural machine translationSeq2seq modelsText summarizationSpeech recognition
How it works
The global Luong variant computes a score between the current decoder state and each encoder state, normalises scores with softmax, and forms the context vector as a weighted sum of encoder states. The local variant first predicts a central source position and then attends only within a window around that position. The score function can be dot, general or concat.
Problem solved
It reduces cost and simplifies attention construction in seq2seq models while also enabling a local variant that restricts the number of source positions considered at each step.
Computational complexity
Time complexity: O(T_x · T_y · d).
Execution paradigm
Primary mode
Dense
Global attention uses all positions; local attention uses a subset of positions.
Activation pattern
All paths active
Additional modes
Conditional
Parallelism
Parallelism level
Sequential
Like Bahdanau, it usually runs inside an RNN decoder, so generation is sequential.
Scope
Inference