Robots Atlas>ROBOTS ATLAS
Evaluation

OSWorld

2024ActivePublished: 1 October 2026Updated: 1 October 2026Published
Key innovation
Moving the evaluation of computer-use agents from simplified, isolated sandboxes into full, real operating systems (Ubuntu, Windows, macOS) with outcome verification based on actual task execution.
Category
Evaluation
Abstraction level
Benchmark
Operation level
Evaluation (runtime)Agent runtime
Use cases
Evaluating computer-use agentsTesting multimodal GUI agentsMeasuring GUI grounding abilityResearch on agents autonomously operating an OSComparing models on real desktop and web tasks

How it works

Each OSWorld task runs in a virtual machine with a real operating system (Ubuntu, Windows or macOS), launched via virtualization/cloud providers (including VMware, VirtualBox, Docker, AWS). At task start the environment is set to a defined initial state, and the agent receives a natural-language instruction together with an observation (a screenshot and optionally an accessibility tree). The agent produces low-level actions (mouse moves and clicks, typing, keyboard shortcuts) that are executed in the machine. On completion a dedicated executable evaluation script runs (the benchmark ships a set of evaluation functions) that inspects the real final state — files, application content, configuration — and assigns a success/failure score. This makes evaluation reproducible and robust against superficial text matching.

Problem solved

Earlier benchmarks for computer-operating agents were limited to narrow, simplified or simulated environments (e.g. a single app or a simplified web page), which failed to capture the complexity of real work in an operating system and made reliable, reproducible evaluation hard. OSWorld addresses this by providing real OS environments with multiple applications and automated, execution-based outcome verification.

Components

Operating system environment (VM)Task execution environment

A virtual machine with a real OS (Ubuntu, Windows, macOS) in which the agent performs actions. Launched via providers such as VMware, VirtualBox, Docker or AWS.

Task suiteEvaluation task set

369 computer tasks built on real web and desktop applications, each with a defined natural-language instruction and initial state.

Execution-based evaluation functionsAutomated, reproducible outcome verification

Scripts that inspect the real final system state after task completion (files, app content, configuration) and assign a success/failure score, instead of text matching.

Implementation

Implementation pitfalls
Unreliability of some evaluation functionsMedium

Some of the original execution-based evaluation functions contained bugs reported by the community, which could under- or over-score results — addressed by the later OSWorld-Verified version.

Fix:Use the OSWorld-Verified version and the latest fixes to the evaluation scripts.

Evolution

Original paper · 2024 · NeurIPS 2024 · Tianbao Xie
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, Tao Yu
2024
OSWorld released (arXiv)
Inflection point

First version of the benchmark with 369 real tasks across Ubuntu, Windows and macOS and execution-based verification.

2024
Presented at NeurIPS 2024

OSWorld is presented at the NeurIPS 2024 conference as a benchmark for computer-use agents.

2025
OSWorld-Verified

An improved version (announced July 28, 2025) fixing community-reported buggy examples and reducing evaluation time with AWS support.