The engine starts from real-world artifacts; coding agents then build a working web environment while browser agents test it; a generate–test–audit–refine loop produces tasks with automatically verifiable success conditions. This yields dozens of environments and thousands of tasks instead of a single fixed suite.
Static, hand-authored web-agent benchmarks saturate quickly and do not scale; WAI automates the production of many environments and tasks with verifiable success, reducing data leakage and enabling continuous evaluation.
A static, hand-built benchmark of realistic web environments (Shuyan Zhou et al., NeurIPS 2024).
Automatic, scalable generation of environments and verifiable tasks (March 2026).
DarwinX reports pass@1 rising to 93.0% (audit-clean) on the official 10-application / 1,260-task suite.