The data engine processes bulk terminal recordings, extracts recurring workflows and turns them into tasks with automatically verifiable success conditions. Tasks are validated, and the best-checked ones form the Verified subset. An agent is scored on whether it correctly completes a multi-step task in a terminal environment.
Hand-curated terminal benchmarks poorly reflect real, complex workflows; TerminalWorld scalably produces authentic tasks from real sessions, complementing (not duplicating) expert-curated benchmarks.
arXiv 2605.22535 (May 21, 2026); a data engine from 80,870 recordings → 1,530 tasks, Verified subset of 200.
DarwinX reports 68.3% (28/41) on its own held-out split of TerminalWorld (Opus 4.8 base).