OSWorld
How it works
Each OSWorld task runs in a virtual machine with a real operating system (Ubuntu, Windows or macOS), launched via virtualization/cloud providers (including VMware, VirtualBox, Docker, AWS). At task start the environment is set to a defined initial state, and the agent receives a natural-language instruction together with an observation (a screenshot and optionally an accessibility tree). The agent produces low-level actions (mouse moves and clicks, typing, keyboard shortcuts) that are executed in the machine. On completion a dedicated executable evaluation script runs (the benchmark ships a set of evaluation functions) that inspects the real final state — files, application content, configuration — and assigns a success/failure score. This makes evaluation reproducible and robust against superficial text matching.
Problem solved
Earlier benchmarks for computer-operating agents were limited to narrow, simplified or simulated environments (e.g. a single app or a simplified web page), which failed to capture the complexity of real work in an operating system and made reliable, reproducible evaluation hard. OSWorld addresses this by providing real OS environments with multiple applications and automated, execution-based outcome verification.
Components
A virtual machine with a real OS (Ubuntu, Windows, macOS) in which the agent performs actions. Launched via providers such as VMware, VirtualBox, Docker or AWS.
369 computer tasks built on real web and desktop applications, each with a defined natural-language instruction and initial state.
Scripts that inspect the real final system state after task completion (files, app content, configuration) and assign a success/failure score, instead of text matching.
Implementation
Some of the original execution-based evaluation functions contained bugs reported by the community, which could under- or over-score results — addressed by the later OSWorld-Verified version.
Evolution
First version of the benchmark with 369 real tasks across Ubuntu, Windows and macOS and execution-based verification.
OSWorld is presented at the NeurIPS 2024 conference as a benchmark for computer-use agents.
An improved version (announced July 28, 2025) fixing community-reported buggy examples and reducing evaluation time with AWS support.