About FLTEval
Definition and scoring
- Organisation
- Mistral AI
- Category
- Coding
- Version
- 2026
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗