A task with its own rules.
A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.
- Organisation
- OpenAI
- Version
- 2026
Multimodal / Benchmark profile
A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.
Observed results
5 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| MiniMax M3 | #1 | minimax-m3 | MiniMax | 91.6 | percent |
| Qwen3.7-Plus | #2 | qwen-3-7-plus | Alibaba Cloud | 91.4 | percent |
| Qwen3.8-27B | #3 | qwen-3-8-27b | Alibaba Cloud | 91.1 | percent |
| Qwen3.6-35B-A3B | #4 | qwen3-6-35b-a3b | Alibaba Cloud | 89.9 | percent |
| Muse Glimmer 30B | #5 | muse-glimmer-30b | Meta | 75.8 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | MiniMax M3MiniMax | Open weights | Source carrier | 91.6% |
| 02 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 91.4% |
| 03 | Qwen3.8-27BAlibaba Cloud | Open weights | Public reference | 91.1% |
| 04 | Qwen3.6-35B-A3BAlibaba Cloud | Open weights | Source carrier | 89.9% |
| 05 | Muse Glimmer 30BMeta | Open weights | Public reference | 75.8% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 91.4% | Relative comparisonprovider-reportedVersion & system1.5 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 91.1% | Public referenceprovider-reportedVersion & system1.5 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.6-27B as published by QwenQwen3.6 27BExact identityqwen3-6-27b-default Canonical product: qwen3-6-27b | 89.4% | Relative comparisonprovider-reportedVersion & system1.5 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Claude Opus 4.6 Max as published by QwenClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 86.6% | Relative comparisonprovider-reportedVersion & system1.5 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64Muse Glimmer 30BExact identitymuse-glimmer-30b-high Canonical product: muse-glimmer-30b | 75.8% | Public referenceprovider-reportedVersion & systemv1.5 / 1,355 prompts Meta modified OmniDocBench v1.5 scoring implementation, four runs | Muse Glimmer Evaluation MethodologyObserved Checked |
Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration.MiniMax M3Exact identitySource label without a registered configuration ID Canonical product: minimax-m3 | 91.6% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 91.4% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration.Qwen3.6-35B-A3BExact identitySource label without a registered configuration ID Canonical product: qwen3-6-35b-a3b | 89.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration.MiniMax M3Exact identitySource label without a registered configuration ID Canonical product: minimax-m3 | 91.6% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 91.4% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
From result to context
A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.