About MLS-Bench Lite
Definition and scoring
- Organisation
- MLS-Bench
- Category
- Coding
- Version
- 2026
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A 30-task subset of MLS-Bench that evaluates whether AI systems can invent generalizable and scalable machine-learning methods. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗