OpenAI · stable
OpenAI's self-hosted 117B-parameter open-weight model with configurable reasoning, a 131,072-token context and maximum output, released under Apache 2.0.
Specialist evidence
Source-native operational evidence for GPT-OSS 120B. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · bash-only SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-oss-120b-high gpt-oss-120b | 26.0% | $0.057 | — | 27.6 calls |
| SWE-bench Owner Leaderboards · owner-current · verified SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-oss-120b-high mini-SWE-agent + gpt-oss-120b | 26.0% | $0.057 | — | 27.6 calls |