Artificial Analysis runs a curated set of evaluations on each model in a standardized way, using pass@1 scoring (share of correct attempts). Individual test scores are weighted within four categories (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%), and the final index is a weighted average across categories. Each test carries its own weight in the index (e.g. GDPval-AA v2 = 20%, Terminal-Bench v2.1 = 16%, τ³-Banking = 14%, HLE = 12%, AA-Omniscience = 12%, SciCode = 8%, AA-LCR = 6%, GPQA Diamond = 6%, CritPt = 6%). Evaluations are repeated many times (>10 repeats), yielding an index confidence interval below ±1%. Version v4.1.1 upgraded the GDPval-AA component from v1 (used in v4.0) to v2, with Elo scoring re-baselined to human expert performance at 1000.
Individual benchmarks measure narrow capabilities and saturate quickly, and each vendor reports a different subset of tests under inconsistent conditions, making direct model comparison hard. The Intelligence Index provides a single, normalized, independently computed value that captures many capability dimensions under consistent conditions.
Agentic evaluation of completing economically valuable work tasks (220 tasks, 44 U.S. occupations); Elo scoring anchored to human experts at 1000.
Agent coordinating knowledge retrieval from unstructured fintech documents and multi-step tool-mediated account modifications (97 tasks).
Terminal-based task execution: software engineering, system administration, data processing, model training, security (89 tasks).
Python programming on scientific computing challenges requiring all unit tests to pass (288 subproblems).
Factual knowledge test penalizing hallucinations across diverse domains (6,000 questions).
Reasoning across multiple long documents (~100k tokens per question), 100 questions.
Frontier academic knowledge: mathematics, humanities, natural sciences (2,158 questions).
Expert scientific knowledge: biology, physics, chemistry (198 hard multiple-choice questions).
Research-level physics reasoning on unpublished problems across physics subfields (70 challenges).
Weighting of the final index across categories: Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%.
Contribution of each test to the index (e.g. GDPval-AA v2 20%, Terminal-Bench v2.1 16%, τ³-Banking 14%).
Evaluations repeated >10 times to keep the index confidence interval below ±1%.
Methodology version (current v4.1.1; v4.0 used GDPval-AA v1; earlier versions: MMLU-Pro, MATH-500, AIME 2025, LiveCodeBench, τ²-Bench Telecom).