The agent receives an instruction phrased the way a law firm partner would phrase it, plus access to a file system mirroring a real working environment: client matter documents, virtual data rooms, templates and background materials. It has tools for working with documents, spreadsheets, presentations and the file system. The output should be a work product ready for partner or client review. Grading rests on an extensive rubric โ over 75,000 expert-written criteria โ but the decisive element is the aggregation rule: the all-pass standard requires every criterion to be met for the task to count. This yields markedly lower scores than the criterion-by-criterion scoring used in BigLaw Bench, but ones closer to the real demands of practice. Harvey publishes the dataset together with an execution harness under an open licence.
Earlier legal benchmarks measured short-horizon reasoning over a single document, which does not reflect work that requires reviewing a bundle of materials, extracting the relevant threads and producing a finished document. LAB moves evaluation to the level of a complete work product and a long task horizon.
Covers transactional, advisory, regulatory and litigation work, embedded in synthetic client matters.
Expert-written criteria, itemized in detail for each task.
A single unmet criterion fails the whole task โ no partial credit.
Matter documents, virtual data rooms, templates and background materials available through file-system tools.
The all-pass standard means even strong models land in the teens or twenties โ comparing those numbers with criterion-by-criterion benchmarks is misleading.
Tasks are embedded in synthetic client matters, which removes confidentiality problems but does not reproduce the full messiness of real files.
In August 2024 Harvey published BigLaw Bench, scoring short legal tasks criterion by criterion.
February 2026 brought the Global variant covering international jurisdictions, and March 2026 the Research variant for case law research.
On 6 May 2026 Harvey released LAB: 1,250+ tasks, 24 practice areas, 75,000 criteria and the all-pass standard, as an open-source project without a leaderboard.