About WeirdML v2
Definition and scoring
- Organisation
- Havard Tveit Ihle
- Category
- Knowledge
- Version
- 2026
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗