AWQ
How it works
AWQ passes a small calibration set through the model and measures activation magnitudes at the input of each linear layer. Weight channels tied to the largest activations are deemed salient. For each channel it derives a scaling factor: the salient channel's weight is multiplied by the scale (reducing its relative quantization error) and the corresponding activation is divided by the same scale, preserving the matmul result. The optimal scale is chosen by grid search that minimizes the layer's output error. After scales are set, weights are quantized group-wise to 4 bits. The method needs no backpropagation or second-order reconstruction, so it is fast and generalizes well across domains.
Problem solved
Naive 4-bit quantization loses accuracy because a small fraction of weights tied to high-activation-magnitude channels is critical to model quality, and uniform quantization damages exactly those channels. AWQ solves this by identifying salient channels from the activation distribution and scaling them so their quantization error is negligible, without labeled training data or backpropagation.
Key mechanisms
Strengths & limitations
Components
Measuring activation magnitudes on a calibration set to identify the weight channels most critical to output quality.
Multiplying a salient channel's weights by a scale and dividing the activation by the same scale, reducing quantization error without changing the matmul result.
After scales are set, weights are quantized group-wise (e.g. groups of 128) to 4-bit integers with a per-group scale.
Implementation
Salient-channel identification relies on activation statistics; a calibration set far from the target data yields suboptimal scales.
AWQ's performance benefit needs kernels that handle packed INT4 weights; running via naive dequantization negates the speedup.
Evolution
The MIT Han Lab presents AWQ as an activation-aware 4-bit quantization method without backpropagation.
The method becomes one of the standards for 4-bit quantization and is integrated into vLLM and TensorRT-LLM.
Hyperparameters (configurable axes)
Target precision of the quantized weights.
Number of weights sharing a scale in group-wise quantization.
Size of the set used to measure activation statistics.
Computational complexity
Time complexity: O(L x G x N_kalib) kalibracja + inferencja jak gęsty model. Space complexity: ~4,x bita / waga (INT4 + skale per kanał).
In the source paper (MLSys 2024, Best Paper) AWQ achieves lower perplexity than round-to-nearest and competes with GPTQ on LLaMA/OPT models at 4-bit, preserving quality on instruction-tuned and multimodal tasks. The quantization process is fast (minutes to hours depending on size) because it needs no training or Hessian inversion.
Compute bottleneck
AWQ's cost is dominated by passing the calibration set through the model and grid-searching scales per layer; inference is memory-bound as in INT4.
Execution paradigm
All channels stay active; only weight scaling is differentiated.
AWQ introduces no routing; it is a weight-quantization method for dense compute.
Parallelism
Calibration proceeds layer by layer (activation statistics) but is parallel within a layer; inference is fully parallel.
Hardware requirements
AWQ models run with optimized INT4 kernels (e.g. in vLLM, TensorRT-LLM, AutoAWQ) on NVIDIA GPUs, combining low memory with high throughput.
AWQ weights can be exported to CPU-runnable formats, though peak performance is on GPUs.