The model is given a task from one of the business domains and a set of tools exposed as REST APIs. It issues a sequence of calls aiming to complete the workflow. The outcome is binary: the task either passes or it does not, with no partial credit and no model-based judge. In parallel the system checks whether guardrails were violated during execution โ that is, whether the agent performed an operation it should not have. In total the set spans 657 tasks, and the aggregated result enters the Artificial Analysis Intelligence Index with a 5% weighting.
Claims about "agents automating business processes" are hard to verify, because success depends on the specific tool environment and on whether the agent breaks organizational rules along the way. AutomationBench-AA provides a reproducible environment with real REST API calls and an explicit record of guardrail violations.
Tasks mirroring SaaS software workflows, spread across different areas of company operations.
The model operates through actual API calls rather than a simulated description of tools.
Independently of task completion, it records whether the agent performed an impermissible operation.
An agent that completed 90% of the workflow and stumbled on the last call scores zero โ results can therefore be jumpy.
AutomationBench-AA is Artificial Analysis's harness instance; its results are not necessarily directly comparable with figures reported on AutomationBench itself.
AutomationBench-AA enters the index with a 5% weighting, covering 657 tasks with REST API tools.