Tasks are built on the O*NET database of work activities: for each of the 44 occupations, industry practitioners compose a set of instructions covering the majority of typical activities. The model must produce a finished work product โ not an answer in a chat window but a file: a document, a spreadsheet, a diagram. Grading rests on pairwise comparisons carried out by experts, who weigh the delivered artefacts for relative quality. In the GDPval-AA v2.1 variant, run by Artificial Analysis as part of its Intelligence Index, models work in an agentic environment with tools and produce file outputs, with 220 agentic tasks graded by blind pairwise judging converted to Elo via a Crowd-BT model; that component carries a 10% weighting in the index.
Academic benchmarks measure reasoning divorced from anything anyone pays for. GDPval ties evaluation to specific occupations and their actual activities, so the result can be interpreted in economic terms rather than merely as a point on an abstract capability scale.
At least 30 tasks per occupation in the full set and 5 in the gold subset; nine US economic sectors.
Tasks must cover the majority of the work activities recorded for the occupation; an occupation counted as knowledge work at 60% or more non-physical activities.
Tasks are composed by industry practitioners, which is meant to ensure realistic instructions and expected work products.
Experts weigh artefacts for relative quality; in the GDPval-AA v2.1 variant blind comparisons are converted to Elo via a Crowd-BT model.
The selection of occupations and sectors rests on the structure of the US economy and the O*NET database; transferring conclusions to other markets requires caution.
OpenAI's GDPval and Artificial Analysis's GDPval-AA v2.1 have different task counts and a different harness โ their results are not directly comparable.
In September 2025 OpenAI published a set of 1,320 tasks from 44 occupations and nine sectors, graded on finished work products.
Artificial Analysis runs its own variant: 220 agentic tasks with file outputs, blind pairwise judging and Elo via a Crowd-BT model, weighted 10% in the index.