Mixed-Granularity Attention
How it works
Two patterns recur across works using this phrase. (1) Parallel fusion: a coarse branch and a fine branch run together and their outputs are merged — e.g. in CFLD multi-scale fine-grained features are injected as bias terms onto a coarse-grained prompt. (2) Hierarchical, adaptive refinement: cheap coarse attention first (over clusters/blocks), then token-level fine attention only where the coarse stage indicates it matters — e.g. Double-P (cluster-level top-p estimation → selective token attention) or Native Sparse Attention (coarse compression + fine-grained token selection). There is no single canonical step-by-step algorithm because there is no single mechanism — the exact steps depend on the paper.
Problem solved
Full self-attention scales quadratically with sequence length, which is costly for long context and high resolution. Purely coarse/global attention, in turn, loses local detail. Mixing granularities aims to reconcile the two goals: cut cost via the coarse level while preserving precision via the fine level.
Components
A cheap attention level over a reduced representation: cluster centroids, pooled/compressed tokens, or region/global summaries.
Official
Precise token- or image-patch-level attention that restores the local precision lost by the coarse level.
Official
Decides where to spend costly fine-grained attention (e.g. top-p thresholding, block selection, MoE-style routing). Present mainly in hierarchical/adaptive variants.
Official
Merges coarse and fine level outputs, e.g. by adding fine features as bias terms (CFLD) or summing selective-attention contributions.
Official
Implementation
The routing/estimation stage (which region to refine) must itself be cheap; otherwise selection cost cancels the savings from the coarse level. Double-P explicitly notes the difficulty of jointly optimizing top-p accuracy, selection overhead, and sparse-attention cost.
Naive gather/scatter over selected tokens is slow without block-structured kernels. NSA is 'hardware-aligned' and UniSparse emphasizes GPU-friendly block-level selection.
Evolution
CFLD (CVPR 2024) proposes a hybrid-granularity attention module encoding multi-scale fine-grained appearance features as bias terms to augment a coarse-grained prompt.
NSA formalizes the hierarchical pattern: coarse-grained token compression plus fine-grained token selection, preserving both global context awareness and local precision; the closest well-established mechanism for LLMs.
Hyperparameters (configurable axes)
Controls how much fine-grained attention to run (e.g. a top-p threshold). Present in adaptive variants such as Double-P.
Granularity of the coarse level: block/cluster size and pooling/downsampling ratio. Governs the cost-quality balance.
How many granularity levels are combined (at least two: coarse and fine; multi-scale works use more).
Window size for the fine/local branch when it uses sliding-window attention.
Computational complexity
Time complexity: Zależna od wariantu: od sub-kwadratowej do O(n²·d).
Execution paradigm
Mode depends on the variant: parallel fusion is dense, hierarchical refinement is sparse/conditional (fine attention only where the coarse stage selects it).
In hierarchical variants the coarse stage governs where fine-grained attention runs. In parallel-fusion variants routing may be absent (both branches always active).