Robots Atlas>ROBOTS ATLAS
Inference

GPTQ

2022ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Post-training quantization based on second-order information (a Hessian approximation) that quantizes weights layer by layer, column by column, correcting the remaining weights after each step to minimize the layer's output reconstruction error — enabling accurate 4- and 3-bit for 100B+ models in a few hours on a single GPU.
Category
Inference
Abstraction level
Pattern
Operation level
ModelPost-trainingInference
Use cases
4-bit and 3-bit quantization of very large LLMs for inferencePreparing models for serving in vLLM / TGI / TensorRT-LLMReducing VRAM while preserving generation qualityQuantizing open-weight models for distribution (e.g. GPTQ formats on Hugging Face)LLM inference on a single consumer GPU

How it works

GPTQ quantizes each linear layer independently, minimizing the reconstruction error of its output on a calibration set. For a given layer it computes the Hessian H = 2 X X^T from the input activations. Weights are quantized column by column: after quantizing each column, the remaining not-yet-quantized weights are updated (error compensation) following a rule derived from Optimal Brain Surgeon, using the inverse Hessian. GPTQ processes columns in a fixed order and applies block updates plus numerical stabilization (Cholesky), letting it quantize billion-parameter models in a few hours on a single GPU. The result is stored as 4- or 3-bit weights with a per-group scale.

Problem solved

Quantizing weights to 4 or 3 bits with round-to-nearest badly degrades the quality of very large models, and methods requiring retraining are too expensive for hundreds-of-billions-parameter models. GPTQ solves this as a fast, one-shot procedure that uses curvature information of the loss (the Hessian of the layer inputs) to quantize weights with minimal increase in output error.

Key mechanisms

Hessian approximation H = 2 X X^T from layer activations
Column-wise quantization with error compensation (OBS rule)
Block updates and a fixed column order
Numerical stabilization: Cholesky and diagonal damping
Storing weights as 4-/3-bit with a per-group scale

Strengths & limitations

Strengths
✓Accurate 4- and 3-bit quantization via second-order information
✓One-shot procedure without retraining
✓Scales to 100B+ models in a few hours on a single GPU
✓Broad ecosystem (AutoGPTQ/GPTQModel, Hugging Face, vLLM, TGI)
✓GPTQ formats as a standard for distributing quantized models
Limitations
✗Numerical instability of an ill-conditioned Hessian
✗Risk of overfitting to a small calibration set
✗Per-layer quantization (local objective), not a global end-to-end error
✗Requires transient memory for the layer Hessian during quantization
✗The performance benefit depends on matched INT4 kernels

Components

Layer Hessian approximationSource of second-order information for error compensation

The matrix H = 2 X X^T from the layer's input activations, describing the output's sensitivity to perturbations of individual weights.

Column-wise quantization with compensationMinimizing the layer output reconstruction error

Quantizing weights column by column, after which the remaining weights are corrected per the OBS rule to reduce accumulated error.

Numerical stabilization (Cholesky)Stability and scalability of the procedure

Cholesky decomposition of the inverse Hessian and block updates ensuring stability and efficiency for large layers.

Implementation

Implementation pitfalls
Hessian numerical instabilityMedium

Inverting the Hessian can be unstable when it is ill-conditioned, corrupting the error compensation.

Fix:Add damping to the Hessian diagonal and use a Cholesky decomposition for a stable solve.
Overfitting to the calibration setMedium

Minimizing reconstruction error on a small, unrepresentative calibration set can hurt the model's generalization.

Fix:Use a sufficiently large and diverse calibration set close to the target data.

Evolution

Original paper · 2022 · ICLR 2023 · Elias Frantar
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh
2022
OBQ / GPTQ formalizes Hessian-based quantization
Inflection point

GPTQ scales the Optimal Brain Quantization method to LLMs, enabling accurate 4- and 3-bit for 100B+ models.

2023
GPTQ formats become a model distribution standard

AutoGPTQ and integration with Hugging Face / Transformers make GPTQ one of the most popular quantized-LLM formats.

Hyperparameters (configurable axes)

Bit widthCritical

Target weight precision (3 or 4 bits).

4The standard with low quality loss.
3More aggressive, greater loss.
Group sizeHigh

Number of weights sharing a scale.

128Typical balance.
Activation order (act-order)Medium

Quantizing columns in order of activation importance, improving accuracy.

trueBetter accuracy, slower kernels.
Hessian dampingMedium

Value added to the Hessian diagonal for numerical stability.

0.01A typical damping value.

Computational complexity

Computational characteristics
→Weight-only quantization, typically group-wise 4- or 3-bit
→Uses second-order information (Hessian)
→Calibration on a few hundred samples, no backpropagation
→Transient memory for the Hessian and its inverse per layer
→Inference with INT4 kernels (ExLlama, Marlin)

Time complexity: O(d^3) na warstwę (odwracanie Hesjanu) + O(N_kalib x d^2). Space complexity: O(d^2) przejściowo na Hesjan + ~4,x bita / waga wynikowo.

Benchmark notes

In the source paper (ICLR 2023) GPTQ quantized OPT-175B and BLOOM-176B to 3-4 bits in about 4 hours on a single A100 GPU, with a small perplexity increase versus FP16. For larger models the 4-bit quality loss is usually marginal, while 3-bit is noticeable but acceptable in many applications.

Compute bottleneck

Layer Hessian operations

During quantization the bottleneck is computing and inverting the Hessian plus sequential error compensation; in inference, as in INT4, memory reads dominate.

Depends on
Szerokość warstwy dUwarunkowanie Hesjanu

Execution paradigm

Primary mode
Dense

All compute paths remain active; the weight representation is quantized.

Activation pattern
All paths active
Routing mechanism

GPTQ introduces no routing; it is a weight-quantization method for dense compute.

Parallelism

Parallelism level
Sequential

Column quantization is inherently sequential (error compensation depends on prior columns); layers can be processed independently, and inference is fully parallel.

Scope
InferenceAcross layers

Hardware requirements

Primary

GPTQ models run with INT4 kernels (ExLlama, Marlin, in vLLM / TGI / TensorRT-LLM) on NVIDIA GPUs, combining low memory with high throughput.

Possible

GPTQ weights can be converted to CPU-runnable formats, though the method and kernels mainly target GPUs.