An agent is given a realistic task (one of 300) and performs it on a live site (one of 136 websites), navigating and interacting like a user. Afterwards, humans judge whether the task was actually completed (human rather than automatic evaluation), giving a reliable success metric. Because the sites are dynamic, the test reflects real conditions (UI, content and availability changes). Results for different agents/models go onto a shared leaderboard.
Older web-agent benchmarks relied on cached pages and automatic scoring, overstating agent success. A measure of real performance on live, changing sites with reliable evaluation was missing.
300 realistic tasks across 136 live websites, reflecting real usage.
Humans judge whether the task was actually completed - instead of automatic, inflating scoring.
Real-time testing against dynamic, evolving interfaces.
Official
A public ranking of agent and model results.
Official
Human evaluation is reliable but costly and slower than automatic scoring.
Changes on live sites can affect reproducibility and comparability of results over time.
OSU-NLP introduced Mind2Web - a benchmark for generalist web agents, based on collected/cached data.
The live, human-evaluated version revealed that agents' real-world success is far below older benchmarks.