About SWE-Lancer
Definition and scoring
- Organisation
- OpenAI
- Category
- Coding
- Version
- Not evaluated
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Low
Freelance-style software engineering benchmark mapping model performance to real paid engineering task outcomes. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗