About BullshitBench v2
Definition and scoring
- Organisation
- Peter Gostev
- Category
- Reasoning
- Version
- 2025
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗