BigCodeBench leaderboard
2 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| DeepSeek V4 Pro Base | #1 | DeepSeek | 59.2% |
| DeepSeek V4 Flash Base | #2 | DeepSeek | 56.8% |
Coding · Benchmark profile
A benchmark for practical code generation involving diverse libraries and complex instructions.
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on BigCodeBench
DeepSeek V4 Pro Base leads at 59.2%, followed by DeepSeek V4 Flash Base (56.8%).
2 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| DeepSeek V4 Pro Base | #1 | DeepSeek | 59.2% |
| DeepSeek V4 Flash Base | #2 | DeepSeek | 56.8% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro Base deepseek-v4-pro-base | DeepSeek | open | Estimated reference | 59.2% |
| #2 | DeepSeek V4 Flash Base deepseek-v4-flash-base | DeepSeek | open | Estimated reference | 56.8% |
About BigCodeBench
A benchmark for practical code generation involving diverse libraries and complex instructions. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A benchmark for practical code generation involving diverse libraries and complex instructions.
DeepSeek V4 Pro Base by DeepSeek currently leads with 59.2%.
2 models in the LuminaBench cohort have a qualifying score on this benchmark.
Related