About Legal Agent Benchmark all-pass rate — Harvey held-out set
Definition and scoring
- Organisation
- Harvey AI
- Category
- Agents
- Version
- 2026
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
Harvey AI's strict held-out task success rate requiring every rubric criterion to pass. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗