Time Window
15/05/2026
01/07/2026
0123456789101112131415161718192021222324252627282930313233343536
Insights
111 problems from 65 repositories selected within the current time window.
111 problems, 65 repositories
Potential contamination
External system
Beyond eval range
#
Model
Resolved Rate (%)
Pass@5 (%)
Cost per Problem ($)
Tokens per Problem
1
Anthropic
Fable 5 [high]
Model
64.5%± 1.41%
78.4%$4.40
2,518,30894.9% cached
2
Grok
Grok 4.5 [high]
Model
63.8%± 0.60%
77.5%$1.47
2,429,42492.6% cached
3
Anthropic
Opus 5 [high]
Model
63.4%± 1.35%
74.8%$3.47
4,322,14395.7% cached
4
Z.ai
GLM-5.2 [high]
Model
62.9%± 1.19%
81.1%$1.40
5,524,89292.1% cached
5
OpenAI
GPT-5.6 Sol [medium]
Model
62.3%± 1.83%
79.3%$0.85
605,34084.7% cached
6
Junie
Junie
Agent
61.8%± 0.54%
73.9%$0.81
1,684,50185.9% cached
7
Anthropic
Claude Code
Agent
60.4%± 1.03%
75.7%$3.39
3,341,58193.4% cached
8
OpenAI
Codex
Agent
58.0%± 1.29%
73.0%$1.59
2,070,97694.6% cached
9
Anthropic
Sonnet 5 [high]
Model
56.8%± 0.94%
74.8%$1.43
4,645,61796.4% cached
10
Cursor
Cursor
Agent
51.7%± 0.84%
65.8%$0.41
1,827,00294.3% cached
11
Minimax
MiniMax M3
Model
47.2%± 1.13%
69.4%$0.95
13,869,45997.0% cached
12
XiaomiMiMo
MiMo V2.5 Pro
Model
46.5%± 0.54%
65.8%$0.10
4,687,98795.7% cached
13
OpenAI
GPT-5.6 Luna [medium]
Model
43.6%± 1.47%
59.5%$0.11
395,52285.2% cached
14
DeepSeek
DeepSeek-V4 Pro [high]
Model
40.2%± 1.29%
64.0%$0.15
3,955,41491.2% cached
15
Qwen
Qwen3.6-27B
Model
31.2%± 1.68%
57.7%$0.62
3,078,97076.8% cached
16
Qwen
Qwen3.6-35B-A3B
Model
24.7%± 0.79%
43.2%$0.27
3,726,06187.5% cached
17
Qwen
Qwen3.5-35B-A3B
Model
17.1%± 1.07%
36.9%$0.99
5,939,22772.3% cached
18
Anthropic
Claude Opus 4.1
Model
N/AN/AN/AN/A
19
Anthropic
Claude Opus 4.5
Model
N/AN/AN/AN/A
20
Anthropic
Claude Opus 4.6-high
Model
N/AN/AN/AN/A
21
Anthropic
Claude Opus 4.7-high
Model
N/AN/AN/AN/A
22
Anthropic
Claude Opus 4.8-xhigh
Model
N/AN/AN/AN/A
23
Anthropic
Claude Sonnet 3.5
Model
N/AN/AN/AN/A
24
Anthropic
Claude Sonnet 4
Model
N/AN/AN/AN/A
25
Anthropic
Claude Sonnet 4.5
Model
N/AN/AN/AN/A
26
Anthropic
Claude Sonnet 4.6
Model
N/AN/AN/AN/A
27
DeepSeek
DeepSeek-R1-0528
Model
N/AN/AN/AN/A
28
DeepSeek
DeepSeek-V3
Model
N/AN/AN/AN/A
29
DeepSeek
DeepSeek-V3-0324
Model
N/AN/AN/AN/A
30
DeepSeek
DeepSeek-V3-0324
Model
N/AN/AN/AN/A
31
DeepSeek
DeepSeek-V3.1
Model
N/AN/AN/AN/A
32
DeepSeek
DeepSeek-V3.2
Model
N/AN/AN/AN/A
33
DeepSeek
DeepSeek-V4 Flash [high]
Model
N/AN/AN/AN/A
34
Mistral
Devstral-2-123B-Instruct-2512
Model
N/AN/AN/AN/A
35
Mistral
Devstral-Small-2-24B-Instruct-2512
Model
N/AN/AN/AN/A
36
Mistral
Devstral-Small-2505
Model
N/AN/AN/AN/A
37
Gemini
Gemini 3 Flash Preview
Model
N/AN/AN/AN/A
38
Gemini
Gemini 3 Pro Preview
Model
N/AN/AN/AN/A
39
Gemini
Gemini 3.1 Pro Preview
Model
N/AN/AN/AN/A
40
Gemini
Gemini 3.5 Flash
Model
N/AN/AN/AN/A
41
Gemini
gemini-2.0-flash
Model
N/AN/AN/AN/A
42
Gemini
gemini-2.0-flash
Model
N/AN/AN/AN/A
43
Gemini
gemini-2.5-flash
Model
N/AN/AN/AN/A
44
Gemini
gemini-2.5-flash-preview-05-20 no-thinking
Model
N/AN/AN/AN/A
45
Gemini
gemini-2.5-flash-preview-05-20 no-thinking
Model
N/AN/AN/AN/A
46
Gemini
gemini-2.5-pro
Model
N/AN/AN/AN/A
47
Gemini
Gemma 4 31B
Model
N/AN/AN/AN/A
48
Gemini
gemma-3-27b-it
Model
N/AN/AN/AN/A
49
Z.ai
GLM-4.5
Model
N/AN/AN/AN/A
50
Z.ai
GLM-4.5 Air
Model
N/AN/AN/AN/A
51
Z.ai
GLM-4.6
Model
N/AN/AN/AN/A
52
Z.ai
GLM-4.7
Model
N/AN/AN/AN/A
53
Z.ai
GLM-4.7 Flash
Model
N/AN/AN/AN/A
54
Z.ai
GLM-5
Model
N/AN/AN/AN/A
55
Z.ai
GLM-5.1
Model
N/AN/AN/AN/A
56
Z.ai
GLM-5.1
Model
N/AN/AN/AN/A
57
OpenAI
gpt-4.1-2025-04-14
Model
N/AN/AN/AN/A
58
OpenAI
gpt-4.1-2025-04-14
Model
N/AN/AN/AN/A
59
OpenAI
gpt-4.1-mini-2025-04-14
Model
N/AN/AN/AN/A
60
OpenAI
gpt-4.1-mini-2025-04-14
Model
N/AN/AN/AN/A
61
OpenAI
gpt-4.1-nano-2025-04-14
Model
N/AN/AN/AN/A
62
OpenAI
gpt-5-2025-08-07-high
Model
N/AN/AN/AN/A
63
OpenAI
gpt-5-2025-08-07-medium
Model
N/AN/AN/AN/A
64
OpenAI
gpt-5-2025-08-07-minimal
Model
N/AN/AN/AN/A
65
OpenAI
gpt-5-codex
Model
N/AN/AN/AN/A
66
OpenAI
gpt-5-mini-2025-08-07-high
Model
N/AN/AN/AN/A
67
OpenAI
gpt-5-mini-2025-08-07-medium
Model
N/AN/AN/AN/A
68
OpenAI
gpt-5.1-codex
Model
N/AN/AN/AN/A
69
OpenAI
gpt-5.1-codex-max
Model
N/AN/AN/AN/A
70
OpenAI
gpt-5.2-2025-12-11-medium
Model
N/AN/AN/AN/A
71
OpenAI
gpt-5.2-2025-12-11-xhigh
Model
N/AN/AN/AN/A
72
OpenAI
gpt-5.2-codex
Model
N/AN/AN/AN/A
73
OpenAI
gpt-5.3-codex
Model
N/AN/AN/AN/A
74
OpenAI
gpt-5.3-codex-xhigh
Model
N/AN/AN/AN/A
75
OpenAI
gpt-5.4-2026-03-05-medium
Model
N/AN/AN/AN/A
76
OpenAI
gpt-5.5-2026-04-23-medium
Model
N/AN/AN/AN/A
77
OpenAI
gpt-5.5-2026-04-23-xhigh
Model
N/AN/AN/AN/A
78
OpenAI
gpt-oss-120b
Model
N/AN/AN/AN/A
79
OpenAI
gpt-oss-120b-high
Model
N/AN/AN/AN/A
80
OpenAI
gpt-oss-20b
Model
N/AN/AN/AN/A
81
Grok
Grok 4
Model
N/AN/AN/AN/A
82
Grok
Grok Code Fast 1
Model
N/AN/AN/AN/A
83
OpenRouter
horizon-alpha
Model
N/AN/AN/AN/A
84
OpenRouter
horizon-beta
Model
N/AN/AN/AN/A
85
Kimi
Kimi K2
Model
N/AN/AN/AN/A
86
Kimi
Kimi K2 Instruct 0905
Model
N/AN/AN/AN/A
87
Kimi
Kimi K2 Thinking
Model
N/AN/AN/AN/A
88
Kimi
Kimi K2.5
Model
N/AN/AN/AN/A
89
Kimi
Kimi K2.6
Model
N/AN/AN/AN/A
90
Meta
Llama-3.3-70B-Instruct
Model
N/AN/AN/AN/A
91
Meta
Llama-4-Maverick-17B-128E-Instruct
Model
N/AN/AN/AN/A
92
Meta
Llama-4-Scout-17B-16E-Instruct
Model
N/AN/AN/AN/A
93
Minimax
MiniMax M2
Model
N/AN/AN/AN/A
94
Minimax
MiniMax M2.1
Model
N/AN/AN/AN/A
95
Minimax
MiniMax M2.5
Model
N/AN/AN/AN/A
96
Minimax
MiniMax M2.7
Model
N/AN/AN/AN/A
97
OpenAI
o3-2025-04-16
Model
N/AN/AN/AN/A
98
OpenAI
o4-mini-2025-04-16
Model
N/AN/AN/AN/A
99
Qwen
Qwen2.5-72B-Instruct
Model
N/AN/AN/AN/A
100
Qwen
Qwen2.5-Coder-32B-Instruct
Model
N/AN/AN/AN/A
101
Qwen
Qwen3-235B-A22B
Model
N/AN/AN/AN/A
102
Qwen
Qwen3-235B-A22B no-thinking
Model
N/AN/AN/AN/A
103
Qwen
Qwen3-235B-A22B thinking
Model
N/AN/AN/AN/A
104
Qwen
Qwen3-235B-A22B-Instruct-2507
Model
N/AN/AN/AN/A
105
Qwen
Qwen3-235B-A22B-Thinking-2507
Model
N/AN/AN/AN/A
106
Qwen
Qwen3-30B-A3B-Instruct-2507
Model
N/AN/AN/AN/A
107
Qwen
Qwen3-30B-A3B-Thinking-2507
Model
N/AN/AN/AN/A
108
Qwen
Qwen3-32B
Model
N/AN/AN/AN/A
109
Qwen
Qwen3-32B no-thinking
Model
N/AN/AN/AN/A
110
Qwen
Qwen3-32B thinking
Model
N/AN/AN/AN/A
111
Qwen
Qwen3-Coder-30B-A3B-Instruct
Model
N/AN/AN/AN/A
112
Qwen
Qwen3-Coder-480B-A35B-Instruct
Model
N/AN/AN/AN/A
113
Qwen
Qwen3-Coder-Next
Model
N/AN/AN/AN/A
114
Qwen
Qwen3-Next-80B-A3B-Instruct
Model
N/AN/AN/AN/A
115
Qwen
Qwen3.5-27B
Model
N/AN/AN/AN/A
116
Qwen
Qwen3.5-397B-A17B
Model
N/AN/AN/AN/A
117
Stepfun
Step-3.5-Flash
Model
N/AN/AN/AN/A

News

  • [2026-07-01]:
    • Added new models to the leaderboad: GLM 5.2, DeepSeek-V4 Pro, DeepSeek-V4 Flash, MiMo V2.5 Pro, Qwen3.6-35B-A3B, Qwen3.6-27B and Gemma 4 31B.
  • [2026-06-09]:
    • Added new models to the leaderboad: Gemini 3.5 Flash and MiniMax M3.
  • [2026-05-28]:
    • Added new models to the leaderboad: Claude Opus 4.8.
  • [2026-05-27]:
    • Added new models to the leaderboad: gpt-5.5-2026-04-23-xhigh, gpt-5.5-2026-04-23-medium, gpt-5.4-2026-03-05-medium, Claude Opus 4.7, and Kimi K2.6.
  • [2026-04-19]:
    • Re-run the Junie with Claude Opus 4.6 as the primary model.
  • [2026-04-15]:
    • Added new models to the leaderboad: GLM-5.1, Qwen3.5-27B, Cursor, Gemma 4 31B and MiniMax M2.7.
  • [2026-03-20]:
    • Added new models to the leaderboard: gpt-5.4-2026-03-05-medium, Gemini 3.1 Pro Preview, Claude Sonnet 4.6, Qwen3.5-397B-A17B, gpt-5.3-codex-xhigh, gpt-5.3-codex and Qwen3.5-35B-A3B
    • Deprecated following models: gpt-5.2-2025-12-11-xhigh, gpt-5.1-codex-max, gpt-5.1-codex, gpt-5-mini-2025-08-07-high, gpt-5-mini-2025-08-07-medium, Qwen3-235B-A22B-Instruct-2507, DeepSeek-R1-0528, Qwen3-Coder-30B-A3B-Instruct, Qwen3-Next-80B-A3B-Instruct and Qwen3-30B-A3B-Instruct-2507.
  • [2026-03-09]:
    • Added reference evaluation for Junie CLI (highlighted in orange). See setup details in Insights.
  • [2026-02-13]:
    • Added new models to the leaderboard: Claude Opus 4.6, GLM-5, MiniMax M2.5, Codex, Qwen3-Coder-Next, GLM-4.7 Flash, gpt-5.2-codex, GLM-4.7 Flash.
  • [2026-01-14]:
    • Added new models to the leaderboard: gpt-5.2-2025-12-11-xhigh, gpt-5.1-codex, GLM-4.7, gpt-5-mini-2025-08-07-high, gpt-oss-120b-high, Kimi K2 Thinking.
    • Deprecated following models: gpt-5-2025-08-07-medium, gpt-5-2025-08-07-high, Claude Sonnet 4, Claude Opus 4.1, o3-2025-04-16, gpt-5-codex, GLM-4.5, o4-mini-2025-04-16, gpt-5-2025-08-07-minimal, gpt-4.1-2025-04-14, Qwen3-235B-A22B-Thinking-2507, gpt-4.1-mini-2025-04-14, Qwen3-30B-A3B-Thinking-2507.
  • [2025-12-22]:
    • Added new model to the leaderboard: MiniMax M2.1.
  • [2025-12-17]:
    • Added new models to the leaderboard: gpt-5.1-codex-max, gpt-5.2-2025-12-11-medium, Devstral-2-123B-Instruct-2512, Devstral-Small-2-24B-Instruct-2512, DeepSeek-V3.2.
    • Added reference evaluation for Claude Code (highlighted in orange). See setup details in Insights.
    • Deprecated following models: gemini-2.5-pro, gemini-2.5-flash, DeepSeek-V3.1.
  • [2025-12-08]:
    • Added new model to the leaderboard: Gemini 3 Pro Preview.
  • [2025-12-05]:
    • Introduced Cached Tokens column.
  • [2025-11-25]:
    • Added new model to the leaderboard: Claude Opus 4.5.
  • [2025-11-13]:
    • Added new model to the leaderboard: MiniMax M2.
  • [2025-10-28]:
    • Added new model to the leaderboard: GLM-4.6.
  • [2025-10-09]:
    • Added new models to the leaderboard: Claude Sonnet 4.5, gpt-5-codex, Claude Opus 4.1, Qwen3-30B-A3B-Thinking-2507 and Qwen3-30B-A3B-Instruct-2507.
    • Added a new Insights section providing analysis and key takeaways from recent model and data releases.
    • Deprecated following models:
      • Text: Llama-3.3-70B-Instruct, Llama-4-Maverick-17B-128E-Instruct, gemma-3-27b-it and Qwen2.5-72B-Instruct.
      • Tools: Claude Sonnet 3.5, Kimi K2, gemini-2.0-flash, Qwen3-235B-A22B and Qwen3-32B.
  • [2025-09-17]:
    • Added new models to the leaderboard: Grok 4, Kimi K2 Instruct 0905, DeepSeek-V3.1 and Qwen3-Next-80B-A3B-Instruct.
  • [2025-09-04]:
    • Added new models to the leaderboard: GLM-4.5, GLM-4.5 Air, Grok Code Fast 1, Kimi K2, gpt-5-mini-2025-08-07-medium, gpt-oss-120b and gpt-oss-20b.
    • Introduced Cost per Problem and Tokens per Problem columns.
    • Added links to the pull requests within the selected time window. You can review them via the Inspect button.
    • Deprecated following models:
      • Text: DeepSeek-V3, DeepSeek-V3-0324, Devstral-Small-2505, gemini-2.0-flash, gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14, Llama-4-Scout-17B-16E-Instruct and Qwen2.5-Coder-32B-Instruct.
      • Tools: horizon-alpha and horizon-beta.
  • [2025-08-12]: Added new models to the leaderboard: gpt-5-medium-2025-08-07, gpt-5-high-2025-08-07 and gpt-5-minimal-2025-08-07.
  • [2025-08-02]: Added new models to the leaderboard: Qwen3-Coder-30B-A3B-Instruct, horizon-beta.
  • [2025-07-31]:
    • Added new models to the leaderboard: gemini-2.5-pro, gemini-2.5-flash, o4-mini-2025-04-16, Qwen3-Coder-480B-A35B-Instruct, Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, DeepSeek-R1-0528 and horizon-alpha.
    • Deprecated models: gemini-2.5-flash-preview-05-20 no-thinking.
    • Updated demo format: tool calls are now shown as distinct assistant and tool messages.
  • [2025-07-11]: Released Docker images for all leaderboard problems and published a dedicated HuggingFace dataset containing only the problems used in the leaderboard.
  • [2025-07-10]: Added models performance chart and evaluations on June data.
  • [2025-06-12]: Added tool usage support, evaluations on May data and new models: Claude Sonnet 3.5/4 and o3.
  • [2025-05-22]: Added Devstral-Small-2505 to the leaderboard.
  • [2025-05-21]: Added new models to the leaderboard: gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14, gemini-2.0-flash and gemini-2.5-flash-preview-05-20.