A task with its own rules.
A benchmark based on challenging real-world user conversations and automated evaluation.
- Organisation
- WildBench authors
- Version
- Current / rolling
Instruction Following / Benchmark profile
A benchmark based on challenging real-world user conversations and automated evaluation.
This benchmark is registered, but the current snapshot has no qualifying results for the model cohort.
Not evaluated for the selected model cohortFrom result to context
A benchmark based on challenging real-world user conversations and automated evaluation.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Archived.