About InferenceBench
Definition and scoring
- Organisation
- Jehyeok Yeon, Ben Rank, Maksym Andriushchenko
- Category
- Agents
- Version
- 2026
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗