About General AI Assistants
Definition and scoring
- Organisation
- BenchLM registry
- Category
- Agents
- Version
- 2024
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗