OpenAI · stable
GPT-4.1 is a proprietary model variant recorded in the BenchLM public dataset.
Specialist evidence
Source-native operational evidence for GPT-4.1. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · bash-only SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-4-1-epoch-gpt-4-1-2025-04-14 GPT 4.1 (2025-04-14) | 39.6% | $0.146 | — | 19.5 calls |
| SWE-bench Owner Leaderboards · owner-current · multimodal SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-4-1-epoch-gpt-4-1-2025-04-14 GUIRepair + GPT 4.1 (2025-04-14) | 31.1% | — | — | — |
| SWE-bench Owner Leaderboards · owner-current · verified SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-4-1-epoch-gpt-4-1-2025-04-14 mini-SWE-agent + GPT 4.1 (2025-04-14) | 39.6% | $0.146 | — | 19.5 calls |