Evaluated with mini-SWE-agent · 200 tasks · Updated Sep. 28, 2026
Rank Model Agent ResolvedRes. help_outline The number of fully solved instances as measured by the hidden behavioral tests. Note that behavioral tests can never cover all possible inputs. The behavioral tests of ProgramBench can be easily extended should any false positives arise. Almost help_outline Instances where the agent's solution solves ≥ 95% of all behavioral tests. Cost help_outline Average API cost in USD per task instance. Calls help_outline Average number of LLM calls per task instance.
1 Claude Opus 5 (xhigh) mini-SWE-agent 4.5% 37.0% $51.04 259
2 Muse Spark 1.3 (max) mini-SWE-agent 2.5% 25.0% $6.46 325
3 Muse Spark 1.3 (xhigh) mini-SWE-agent 1.0% 16.5% $2.02 253
4 GPT-5.6 Sol (xhigh) mini-SWE-agent 1.0% 15.5% $6.08 42
5 GPT 5.5 (xhigh) mini-SWE-agent 0.5% 13.5% $8.85 82
6 GPT 5.5 (high) mini-SWE-agent 0.5% 5.0% $3.65 41
7 Gemini 3.6 Flash mini-SWE-agent 0.5% 4.0% $4.88 151
8 GPT-5.6 Sol mini-SWE-agent 0.5% 2.5% $1.02 15
9 Claude Opus 4.8 (xhigh) mini-SWE-agent 0% 16.5% $21.02 145
10 GLM-5.2 mini-SWE-agent 0% 8.5% $25.49 223
11 Gemini 3.7 Flash mini-SWE-agent 0% 5.5% $2.05 134
12 Muse Spark 1.2 (xhigh) mini-SWE-agent 0% 4.5% $1.50 185
13 Claude Opus 4.7 (xhigh) mini-SWE-agent 0% 4.5% $10.96 159
14 Muse Spark 1.1 (xhigh) mini-SWE-agent 0% 4.0% $0.73 195
15 Gemini 3.5 Flash mini-SWE-agent 0% 3.0% $5.66 168
16 Claude Opus 4.7 mini-SWE-agent 0% 3.0% $3.81 93
17 Claude Opus 4.6 mini-SWE-agent 0% 2.5% $11.38 260
18 GPT 5.5 mini-SWE-agent 0% 1.5% $1.21 19
19 Claude Sonnet 4.6 mini-SWE-agent 0% 1.0% $26.73 472
20 GPT 5.4 mini-SWE-agent 0% 0.0% $0.33 16
21 Gemini 3.1 Pro mini-SWE-agent 0% 0.0% $1.51 94
22 Gemini 3 Flash mini-SWE-agent 0% 0.0% $0.30 85
23 Claude Haiku 4.5 mini-SWE-agent 0% 0.0% $0.80 124
24 GPT 5.4 mini mini-SWE-agent 0% 0.0% $0.04 18
25 GPT 5 mini mini-SWE-agent 0% 0.0% $0.03 15

Click row to see model details · Sorting: Resolved → Almost resolved → Avg. pass rate (more)

Hover for details · Click a model to open it · Drag to zoom the x-axis · Shift-drag to pan · The line marks the Pareto frontier (best score per cost)

All models × 200 tasks
0%
100%

Hover for details · Click to open task

Double click legend item (show only this model) · Click (hide model)

Each dot is one task instance · Hover for details · Click to view task