OpenAI · stable
o4-mini is tracked via Artificial Analysis independent evaluations for capability comparison across coding, agents, reasoning and related workloads.
Specialist evidence
Source-native operational evidence for o4-mini. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · bash-only SWE-bench Owner Leaderboards · checked 2026-08-29 | o4-mini-default o4-mini (2025-04-16) | 45.0% | $0.210 | — | 23.3 calls |
| SWE-bench Owner Leaderboards · owner-current · multimodal SWE-bench Owner Leaderboards · checked 2026-08-29 | o4-mini-default GUIRepair + o4-mini (2025-04-16) | 33.9% | — | — | — |
| SWE-bench Owner Leaderboards · owner-current · verified SWE-bench Owner Leaderboards · checked 2026-08-29 | o4-mini-default mini-SWE-agent + o4-mini (2025-04-16) | 45.0% | $0.210 | — | 23.3 calls |