GPTQ
How it works
GPTQ quantizes each linear layer independently, minimizing the reconstruction error of its output on a calibration set. For a given layer it computes the Hessian H = 2 X X^T from the input activations. Weights are quantized column by column: after quantizing each column, the remaining not-yet-quantized weights are updated (error compensation) following a rule derived from Optimal Brain Surgeon, using the inverse Hessian. GPTQ processes columns in a fixed order and applies block updates plus numerical stabilization (Cholesky), letting it quantize billion-parameter models in a few hours on a single GPU. The result is stored as 4- or 3-bit weights with a per-group scale.
Problem solved
Quantizing weights to 4 or 3 bits with round-to-nearest badly degrades the quality of very large models, and methods requiring retraining are too expensive for hundreds-of-billions-parameter models. GPTQ solves this as a fast, one-shot procedure that uses curvature information of the loss (the Hessian of the layer inputs) to quantize weights with minimal increase in output error.
Key mechanisms
Strengths & limitations
Components
The matrix H = 2 X X^T from the layer's input activations, describing the output's sensitivity to perturbations of individual weights.
Quantizing weights column by column, after which the remaining weights are corrected per the OBS rule to reduce accumulated error.
Cholesky decomposition of the inverse Hessian and block updates ensuring stability and efficiency for large layers.
Implementation
Inverting the Hessian can be unstable when it is ill-conditioned, corrupting the error compensation.
Minimizing reconstruction error on a small, unrepresentative calibration set can hurt the model's generalization.
Evolution
GPTQ scales the Optimal Brain Quantization method to LLMs, enabling accurate 4- and 3-bit for 100B+ models.
AutoGPTQ and integration with Hugging Face / Transformers make GPTQ one of the most popular quantized-LLM formats.
Hyperparameters (configurable axes)
Target weight precision (3 or 4 bits).
Number of weights sharing a scale.
Quantizing columns in order of activation importance, improving accuracy.
Value added to the Hessian diagonal for numerical stability.
Computational complexity
Time complexity: O(d^3) na warstwę (odwracanie Hesjanu) + O(N_kalib x d^2). Space complexity: O(d^2) przejściowo na Hesjan + ~4,x bita / waga wynikowo.
In the source paper (ICLR 2023) GPTQ quantized OPT-175B and BLOOM-176B to 3-4 bits in about 4 hours on a single A100 GPU, with a small perplexity increase versus FP16. For larger models the 4-bit quality loss is usually marginal, while 3-bit is noticeable but acceptable in many applications.
Compute bottleneck
During quantization the bottleneck is computing and inverting the Hessian plus sequential error compensation; in inference, as in INT4, memory reads dominate.
Execution paradigm
All compute paths remain active; the weight representation is quantized.
GPTQ introduces no routing; it is a weight-quantization method for dense compute.
Parallelism
Column quantization is inherently sequential (error compensation depends on prior columns); layers can be processed independently, and inference is fully parallel.
Hardware requirements
GPTQ models run with INT4 kernels (ExLlama, Marlin, in vLLM / TGI / TensorRT-LLM) on NVIDIA GPUs, combining low memory with high throughput.
GPTQ weights can be converted to CPU-runnable formats, though the method and kernels mainly target GPUs.