Robots Atlas>ROBOTS ATLAS
Multimodal

UGS

2026ResearchPublished
Key innovation
Combines visual grounding and retrieval into a single atomic action handling N entities concurrently, replacing sequential tool calls with a single parallel query.
Category
Multimodal
Abstraction level
Building block
Operation level
Architecture blockInferenceAgent runtime
Use cases
Visual search in e-commerce (multiple products in one image)Multi-entity document analysis (people, places, dates)Q&A agents for complex imagesMultimodal information retrieval systemsSocial platform content analysis (e.g. XiaoHongShu)

How it works

For a given query (image + text question), the model identifies all entities requiring search, simultaneously generates bounding boxes (visual grounding) and retrieval queries for them in a single atomic action. Results from parallel searches are aggregated and the model generates the final answer. Example: question about 6 people in a photo โ†’ 1 UGS action โ†’ 6 parallel searches โ†’ answer in 3 rounds instead of 12.

Problem solved

Sequential multimodal agents process one entity per round โ€” for queries with N entities this generates N tool-call rounds, accumulating latency, token costs, and error propagation risk. UGS eliminates this bottleneck.

Implementation

Reference implementations
Implementation pitfalls
Overly broad grounding reduces precisionMedium

When the model attempts to ground too many entities simultaneously, bounding boxes may overlap or cover incorrect image regions, degrading retrieval quality.

Dependency on base model visual grounding qualityMedium

UGS is only as good as the base model's visual grounding โ€” poor grounding on complex images (crowds, small objects) directly results in incorrect retrieval queries.

Parallel tool calls increase cost per roundMedium

A single UGS action triggers N parallel tool calls โ€” for queries with many entities the cost per round is higher than in a sequential agent, even though the total number of rounds is lower.

Execution paradigm

Primary mode
Sparse
Activation pattern
Input dependent

Parallelism

Parallelism level
Fully parallel
Scope
Inference

Hardware requirements

Parallel visual grounding and retrieval processing for N entities requires GPU for efficient multimodal model inference.