Robots Atlas>ROBOTS ATLAS
Architecture

Mixed-Granularity Attention

ActivePublished: 1 October 2026Updated: 1 October 2026Published
Key innovation
Computing attention at several granularity levels at once — coarse (clusters/regions/pooled or global summaries) and fine (token/patch level) — instead of a single uniform granularity, to reconcile compute cost with preserving local detail.
Category
Architecture
Abstraction level
Pattern
Operation level
Architecture blockLayer
Use cases
Sparse/efficient attention for long-context LLMs (coarse selection + fine-grained attention)Pose-guided image generation (injecting multi-scale fine-grained features)Multi-scale representation in computer vision (point clouds, person re-ID)Reducing attention cost while preserving local precision

How it works

Two patterns recur across works using this phrase. (1) Parallel fusion: a coarse branch and a fine branch run together and their outputs are merged — e.g. in CFLD multi-scale fine-grained features are injected as bias terms onto a coarse-grained prompt. (2) Hierarchical, adaptive refinement: cheap coarse attention first (over clusters/blocks), then token-level fine attention only where the coarse stage indicates it matters — e.g. Double-P (cluster-level top-p estimation → selective token attention) or Native Sparse Attention (coarse compression + fine-grained token selection). There is no single canonical step-by-step algorithm because there is no single mechanism — the exact steps depend on the paper.

Problem solved

Full self-attention scales quadratically with sequence length, which is costly for long context and high resolution. Purely coarse/global attention, in turn, loses local detail. Mixing granularities aims to reconcile the two goals: cut cost via the coarse level while preserving precision via the fine level.

Components

Coarse-grained branchCheap, approximate capture of global context

A cheap attention level over a reduced representation: cluster centroids, pooled/compressed tokens, or region/global summaries.

Official

Fine-grained branchPreserves local detail and precision

Precise token- or image-patch-level attention that restores the local precision lost by the coarse level.

Official

Router / selection mechanismAllocates compute budget across granularity levels

Decides where to spend costly fine-grained attention (e.g. top-p thresholding, block selection, MoE-style routing). Present mainly in hierarchical/adaptive variants.

Official

Fusion stepCombines multi-granularity outputs into one representation

Merges coarse and fine level outputs, e.g. by adding fine features as bias terms (CFLD) or summing selective-attention contributions.

Official

Implementation

Implementation pitfalls
Selection overhead can dominate the gainsHigh

The routing/estimation stage (which region to refine) must itself be cheap; otherwise selection cost cancels the savings from the coarse level. Double-P explicitly notes the difficulty of jointly optimizing top-p accuracy, selection overhead, and sparse-attention cost.

Fix:Use cheap, block-wise estimation and block structure instead of per-token selection.
Hardware alignment is criticalMedium

Naive gather/scatter over selected tokens is slow without block-structured kernels. NSA is 'hardware-aligned' and UniSparse emphasizes GPU-friendly block-level selection.

Fix:Design selection and kernels in a block layout aligned with GPU architecture.

Evolution

2024
CFLD introduces 'hybrid-granularity attention' (closest named relative)
Inflection point

CFLD (CVPR 2024) proposes a hybrid-granularity attention module encoding multi-scale fine-grained appearance features as bias terms to augment a coarse-grained prompt.

2025
Native Sparse Attention (NSA): canonical coarse-compression + fine-grained selection

NSA formalizes the hierarchical pattern: coarse-grained token compression plus fine-grained token selection, preserving both global context awareness and local precision; the closest well-established mechanism for LLMs.

Hyperparameters (configurable axes)

Selection threshold / attention-mass budgetHigh

Controls how much fine-grained attention to run (e.g. a top-p threshold). Present in adaptive variants such as Double-P.

Cluster/block size and compression ratioHigh

Granularity of the coarse level: block/cluster size and pooling/downsampling ratio. Governs the cost-quality balance.

Number of granularity levels / scalesMedium

How many granularity levels are combined (at least two: coarse and fine; multi-scale works use more).

Local window sizeMedium

Window size for the fine/local branch when it uses sliding-window attention.

Computational complexity

Time complexity: Zależna od wariantu: od sub-kwadratowej do O(n²·d).

Execution paradigm

Primary mode
Sparse

Mode depends on the variant: parallel fusion is dense, hierarchical refinement is sparse/conditional (fine attention only where the coarse stage selects it).

Activation pattern
Input dependent
Additional modes
DenseConditional
Routing mechanism

In hierarchical variants the coarse stage governs where fine-grained attention runs. In parallel-fusion variants routing may be absent (both branches always active).