The model receives a task mirroring part of a multi-week business project and must produce output files. Grading combines two mechanisms. The rubric checks verifiable task success โ whether what was asked for was delivered in the required form. Pairwise comparisons assess analytical quality and presentation quality: outputs from different models are set against each other in pairs, and the resulting order is converted into Elo ratings. The three components โ analytical quality, presentation quality and rubric score โ are then aggregated. The v1.1 set contains 91 tasks and enters the Artificial Analysis Intelligence Index with a 15% weighting.
Classic benchmarks measure single answers, whereas knowledge work consists of producing a document, spreadsheet or presentation whose quality cannot be reduced to one correct string. AA Briefcase evaluates real artefacts and separates what can be verified objectively from what requires qualitative comparison.
The tasks mirror realistic knowledge work spread across multi-week business projects.
Checks verifiable task success โ whether the required elements were delivered in the expected form.
Analytical and presentation quality assessed by pitting outputs against each other in pairs, converted into Elo ratings.
Elo ratings are only meaningful relative to the specific field of compared models; adding or removing a participant shifts the scale.
Part of the score concerns presentation quality, which can reward visually polished outputs over substantively sharper but rawer ones.
AA Briefcase v1.1 covers 91 tasks with file outputs and becomes the highest-weighted component of the index, level with AA-Omniscience.