Robots Atlas>ROBOTS ATLAS
Inference

AWQ

2023ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
Activation-aware weight quantization: instead of treating all weights equally, AWQ protects the ~1% most important weight channels (identified from the activation distribution) via per-channel scaling, achieving accurate 4-bit quantization without backpropagation.
Category
Inference
Abstraction level
Pattern
Operation level
ModelPost-trainingInference
Use cases
4-bit LLM quantization for local and server inferenceDeploying models on edge and mobile devicesPreparing models for serving in vLLM / TensorRT-LLMQuantizing multimodal and instruction-tuned models while preserving qualityReducing VRAM while keeping high inference throughput

How it works

AWQ passes a small calibration set through the model and measures activation magnitudes at the input of each linear layer. Weight channels tied to the largest activations are deemed salient. For each channel it derives a scaling factor: the salient channel's weight is multiplied by the scale (reducing its relative quantization error) and the corresponding activation is divided by the same scale, preserving the matmul result. The optimal scale is chosen by grid search that minimizes the layer's output error. After scales are set, weights are quantized group-wise to 4 bits. The method needs no backpropagation or second-order reconstruction, so it is fast and generalizes well across domains.

Problem solved

Naive 4-bit quantization loses accuracy because a small fraction of weights tied to high-activation-magnitude channels is critical to model quality, and uniform quantization damages exactly those channels. AWQ solves this by identifying salient channels from the activation distribution and scaling them so their quantization error is negligible, without labeled training data or backpropagation.

Key mechanisms

Analysis of activation magnitudes on a calibration set
Identification of the ~1% salient weight channels
Per-channel scaling (weight x s, activation / s) that preserves the result
Grid search over scales minimizing the layer output error
Group-wise 4-bit quantization after scales are set

Strengths & limitations

Strengths
✓Accurate 4-bit quantization preserving model quality
✓No backpropagation or second-order reconstruction (speed)
✓Good generalization across domains and instruction-tuned/multimodal models
✓Low risk of overfitting to the calibration set
✓Optimized kernel support in vLLM, TensorRT-LLM, AutoAWQ
Limitations
✗Requires a representative calibration set for activation statistics
✗The performance benefit depends on matched INT4 kernels
✗Focused on weight (not activation) quantization - mostly weight-only
✗Extremely low bit widths (2-3) remain hard without quality loss
✗Grid-search scale selection adds a calibration step before deployment

Components

Activation saliency analysisSelecting salient channels to protect

Measuring activation magnitudes on a calibration set to identify the weight channels most critical to output quality.

Per-channel scalingProtecting salient weights from quantization error

Multiplying a salient channel's weights by a scale and dividing the activation by the same scale, reducing quantization error without changing the matmul result.

Group-wise 4-bit quantizationThe actual compression of weights to 4 bits

After scales are set, weights are quantized group-wise (e.g. groups of 128) to 4-bit integers with a per-group scale.

Implementation

Implementation pitfalls
Unrepresentative calibration setMedium

Salient-channel identification relies on activation statistics; a calibration set far from the target data yields suboptimal scales.

Fix:Use a few hundred samples representative of the target domain; AWQ is robust but not to a wildly mismatched distribution.
Missing matched inference kernelsMedium

AWQ's performance benefit needs kernels that handle packed INT4 weights; running via naive dequantization negates the speedup.

Fix:Serve through an AWQ-aware runtime (vLLM, TensorRT-LLM, AutoAWQ) with native INT4 kernels.

Evolution

Original paper · 2023 · MLSys 2024 (Best Paper) · Ji Lin
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Guangxuan Xiao, Song Han
2023
AWQ published
Inflection point

The MIT Han Lab presents AWQ as an activation-aware 4-bit quantization method without backpropagation.

2024
AWQ wins Best Paper at MLSys 2024

The method becomes one of the standards for 4-bit quantization and is integrated into vLLM and TensorRT-LLM.

Hyperparameters (configurable axes)

Weight bit widthCritical

Target precision of the quantized weights.

4The most common configuration.
3More aggressive, greater quality loss.
Group sizeHigh

Number of weights sharing a scale in group-wise quantization.

128Default balance.
Calibration samplesMedium

Size of the set used to measure activation statistics.

128-512Typical range.

Computational complexity

Computational characteristics
→Weight-only quantization, typically group-wise 4-bit
→Calibration on a few hundred samples, no backpropagation
→Per-channel scale for salient weights
→Resulting model ~4x smaller than FP16
→Inference with INT4 kernels (vLLM, TensorRT-LLM, AutoAWQ)

Time complexity: O(L x G x N_kalib) kalibracja + inferencja jak gęsty model. Space complexity: ~4,x bita / waga (INT4 + skale per kanał).

Benchmark notes

In the source paper (MLSys 2024, Best Paper) AWQ achieves lower perplexity than round-to-nearest and competes with GPTQ on LLaMA/OPT models at 4-bit, preserving quality on instruction-tuned and multimodal tasks. The quantization process is fast (minutes to hours depending on size) because it needs no training or Hessian inversion.

Compute bottleneck

Activation calibration and scale search

AWQ's cost is dominated by passing the calibration set through the model and grid-searching scales per layer; inference is memory-bound as in INT4.

Depends on
Rozmiar zbioru kalibracyjnegoGęstość siatki skal

Execution paradigm

Primary mode
Dense

All channels stay active; only weight scaling is differentiated.

Activation pattern
All paths active
Routing mechanism

AWQ introduces no routing; it is a weight-quantization method for dense compute.

Parallelism

Parallelism level
Partially parallel

Calibration proceeds layer by layer (activation statistics) but is parallel within a layer; inference is fully parallel.

Scope
InferenceAcross layers

Hardware requirements

Primary

AWQ models run with optimized INT4 kernels (e.g. in vLLM, TensorRT-LLM, AutoAWQ) on NVIDIA GPUs, combining low memory with high throughput.

Possible

AWQ weights can be exported to CPU-runnable formats, though peak performance is on GPUs.