OpenAI · stable
GPT-5 is a API model tracked via Artificial Analysis independent evaluations for capability comparison across coding, agents, and reasoning.
Specialist evidence
Source-native operational evidence for GPT-5. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · bash-only SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-5-medium GPT 5 (2025-08-07) (medium) | 65.0% | $0.280 | — | 13.2 calls |
| SWE-bench Owner Leaderboards · owner-current · verified SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-5-default OpenHands + GPT-5 | 71.8% | — | — | — |
| SWE-bench Owner Leaderboards · owner-current · verified SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-5-medium mini-SWE-agent + GPT 5 (2025-08-07) (medium) | 65.0% | $0.280 | — | 13.2 calls |