Evidence, not our own benchmarks

Every number below was run by somebody else and published openly. We did not run any of it. What CacheSphere adds is the comparison.

Sources: the SWE-bench project's own experiments repository. Each model links to the exact submission directory its figures come from, so you can audit any row.

What the data says

On the same harness (mini-v2.0.0), Claude 4.5 Opus resolved the most real issues — 384 of the 500-instance suite, or 76.8%. Minimax 2.5 came within 1.0 points of that at $0.073 per task against $0.75 — roughly 10× cheaper.

Two columns matter. Resolve rate is always measured against the whole split, so a submission that attempted fewer instances still counts every instance it skipped. Cost per task is the spend the submitting harness reported, divided by instances attempted. The expensive model wins on rate; the cheap one wins on almost everything else.

Every submission we found

256 submissions across 179 distinct models, 4 benchmark suites, from 2 independent publishers. Every row links to the exact submission it came from.

Aider polyglot 64 submissions

ResolvedModelInstancesCost / taskHarnessRun
88.0% gpt-5 225 — diff 2025-08-23
84.9% o3-pro 225 — diff 2025-06-28
83.1% gemini-2.5-pro-preview-06-05 (32k think) 225 — diff-fenced 2025-06-06
81.3% o3 225 — diff 2025-06-25
79.6% grok-4 225 — diff 2025-07-11
79.1% gemini-2.5-pro-preview-06-05 (default think) 225 — diff-fenced 2025-06-06
78.2% o3 (high) + gpt-4.1 225 — architect 2025-06-27
76.9% Gemini 2.5 Pro Preview 05-06 225 — diff-fenced 2025-05-07
74.2% DeepSeek-V3.2-Exp (Reasoner) 225 — diff 2025-10-03
72.9% Gemini 2.5 Pro Preview 03-25 225 — diff-fenced 2025-04-12
72.0% o4-mini 225 — diff 2025-04-16
72.0% claude-opus-4-20250514 (32k thinking) 225 — diff 2025-05-25
71.4% DeepSeek R1 (0528) 225 — diff 2025-06-06
70.7% claude-opus-4-20250514 (no think) 225 — diff 2025-05-25
70.2% DeepSeek-V3.2-Exp (Chat) 225 — diff 2025-10-03
64.9% claude-3-7-sonnet-20250219 (32k thinking tokens) 225 — diff 2025-02-24
64.0% DeepSeek R1 + claude-3-5-sonnet-20241022 225 — architect 2025-01-23
61.7% o1-2024-12-17 225 — diff 2024-12-21
61.3% claude-sonnet-4-20250514 (32k thinking) 225 — diff 2025-05-24
60.4% o3-mini 225 — diff 2025-01-31
60.4% claude-3-7-sonnet-20250219 (no thinking) 225 — diff 2025-02-24
59.6% Qwen3 235B A22B diff, no think, Alibaba API 225 — diff 2025-05-09
59.1% Kimi K2 225 — diff 2025-07-17
56.9% DeepSeek R1 225 — diff 2025-01-20
56.4% claude-sonnet-4-20250514 (no thinking) 225 — diff 2025-05-24
55.1% DeepSeek V3 (0324) 225 — diff 2025-03-24
55.1% gemini-2.5-flash-preview-05-20 (24k think) 225 — diff 2025-05-25
54.7% Quasar Alpha 225 — diff 2025-04-04
53.3% Grok 3 Beta 225 — diff 2025-04-10
52.9% Optimus Alpha 225 — diff 2025-04-10
52.4% gpt-4.1 225 — diff 2025-04-14
51.6% claude-3-5-sonnet-20241022 225 — diff 2025-01-17
49.3% Grok 3 Mini Beta 225 — whole 2025-04-10
48.4% DeepSeek Chat V3 (prev) 225 — diff 2024-12-25
47.1% gemini-2.5-flash-preview-04-17 225 — diff 2025-04-20
45.3% chatgpt-4o-latest (2025-03-29) 225 — diff 2025-03-29
44.9% gpt-4.5-preview 225 — diff 2025-02-27
44.0% gemini-2.5-flash-preview-05-20 (no think) 225 — diff 2025-05-26
41.8% gpt-oss-120b 225 — diff 2025-08-06
40.0% Qwen3 32B 225 — diff 2025-05-08
38.2% gemini-exp-1206 225 — whole 2024-12-22
35.6% Gemini 2.0 Pro exp-02-05 225 — whole 2025-02-25
32.9% o1-mini-2024-09-12 225 — whole 2024-12-22
32.4% gpt-4.1-mini 225 — diff 2025-04-14
28.0% claude-3-5-haiku-20241022 225 — diff 2024-12-21
27.1% chatgpt-4o-latest (2025-02-15) 225 — diff 2025-02-15
26.2% QwQ-32B + Qwen 2.5 Coder Instruct 225 — architect 2025-03-07
23.1% gpt-4o-2024-08-06 225 — diff 2024-12-30
22.2% gemini-2.0-flash-exp 225 — whole 2024-12-22
21.8% qwen-max-2025-01-25 225 — diff 2025-01-28
20.9% QwQ-32B 225 — diff 2025-03-06
18.2% gpt-4o-2024-11-20 225 — diff 2024-12-30
18.2% gemini-2.0-flash-thinking-exp-01-21 225 — diff 2025-01-21
17.8% DeepSeek Chat V2.5 225 — diff 2024-12-21
16.4% Qwen2.5-Coder-32B-Instruct 225 — whole 2024-12-26
15.6% Llama 4 Maverick 225 — whole 2025-04-06
12.9% yi-lightning 225 — whole 2024-12-23
12.0% command-a-03-2025-quality 225 — whole 2025-03-14
11.1% Codestral 25.01 225 — whole 2025-01-13
10.2% openhands-lm-32b-v0.1 225 — whole 2025-04-19
8.9% gpt-4.1-nano 225 — whole 2025-04-14
8.0% Qwen2.5-Coder-32B-Instruct 225 — diff 2024-12-22
4.9% gemma-3-27b-it 225 — whole 2025-03-15
3.6% gpt-4o-mini-2024-07-18 225 — whole 2024-12-21

SWE-bench Lite 59 submissions

ResolvedModelInstancesCost / taskHarnessRun
51.3% Claude 3.5 Sonnet 20241022 154/300 (154 tried) — isea 2025-09-11
58.3% Claude 4 Sonnet 20250514 175/300 (179 tried) — KGCompass 2025-09-06
49.7% R2E_QwenCoder30BA3B_tts 149/300 (149 tried) — entroPO 2025-09-01
45.0% R2E_QwenCoder30BA3B 135/300 (137 tried) — entroPO 2025-09-01
16.3% MCTS Refine 7B 49/300 (87 tried) — agentless 2025-06-27
60.3% Claude 4 Sonnet 20250514 181/300 (181 tried) — ExpeRepair-v1 2025-06-25
51.7% Multi V1_Claude3.7Sonnet_Gemini2.5Pro 155/300 (155 tried) — SemAgent 2025-06-25
46.0% Claude 3.5 Sonnet 20241022 138/300 (147 tried) — KGCompass 2025-06-19
36.7% DeepSeek V3 110/300 (116 tried) — KGCompass 2025-06-09
56.7% Claude 4 Sonnet 20250514 170/300 (171 tried) — sweagent 2025-05-26
42.7% Claude 3.5 Sonnet 20241022 128/300 (129 tried) — Lingxi 2025-05-09
60.0% Agent 180/300 (181 tried) — Refact 2025-04-25
24.7% Qwen2.5 7b Retriever_Qwen2.5 72b Editor 74/300 (78 tried) — SWE-Fixer 2025-03-06
48.0% Claude 3.7 Sonnet 20250219 144/300 (146 tried) — sweagent 2025-02-26
32.3% Lite_o3_mini 97/300 (103 tried) — agentless 2025-02-14
30.3% O3mini 91/300 (130 tried) — aegis 2025-02-07
47.0% Agent_claude_3.5_sonnet_deepseek_r1 141/300 (141 tried) — dars 2025-02-05
39.0% Claude 3.5 Sonnet 20241022 117/300 (128 tried) — moatless 2025-01-14
39.7% Gpt4o 119/300 (133 tried) — OpenCSG-Starship-Agentic-Coder 2025-01-13
30.7% Deepseek_v3 92/300 (105 tried) — moatless 2025-01-11
37.0% Codes_claude 3.5 Sonnet 20241022 111/300 (224 tried) — patched 2025-01-04
49.0% Agent_v1 147/300 (159 tried) — blackboxai 2024-12-20
41.3% Claude 3.5 Sonnet 20241022 124/300 (129 tried) — PatchKitty-0.9 2024-12-20
44.7% Sonnet_v1 134/300 (208 tried) — kodu 2024-12-07
40.7% Claude 3.5 Sonnet 20241022 122/300 (124 tried) — agentless-1.5 2024-12-02
23.3% Qwen2.5 7b Retriever_Qwen2.5 72b Editor_20241128 70/300 (70 tried) — SWE-Fixer 2024-11-28
48.3% Codefixer_agent 145/300 (160 tried) — globant 2024-11-27
38.3% Claude 3.5 Sonnet 20241022 115/300 (126 tried) — moatless 2024-11-17
28.0% Gpt4o 84/300 (91 tried) — reproducedRG 2024-11-17
31.3% Gpt4o 94/300 (106 tried) — codeshelltester 2024-11-11
41.0% Swekit 123/300 (124 tried) — composio 2024-10-30
32.0% Gpt4o 96/300 (100 tried) — agentless-1.5 2024-10-28
25.3% Lite1 76/300 (119 tried) — hyperagent 2024-09-25
30.0% Gpt4o 90/300 (117 tried) — infant 2024-09-08
21.7% Mixed 65/300 (96 tried) — autose 2024-08-28
29.7% Gpt4o 89/300 (96 tried) — RepoGraph 2024-08-08
18.3% Gpt4o 55/300 (94 tried) — sweagent 2024-07-28
26.7% Codeact_v1.8_claude35sonnet 80/300 (113 tried) — opendevin 2024-07-25
27.7% Gpt4o 83/300 — sima 2024-07-06
43.0% Aide_mixed 129/300 (273 tried) — codestory 2024-07-02
27.3% Gpt4o 82/300 (292 tried) — agentless 2024-06-30
38.0% Mentatbot_gpt4o 114/300 (296 tried) — abanteai 2024-06-27
26.7% Claude35sonnet 80/300 (294 tried) — moatless 2024-06-23
33.0% Agent 99/300 (292 tried) — Lingma 2024-06-22
23.0% Claude3.5sonnet 69/300 (81 tried) — sweagent 2024-06-20
31.3% Code_droid 94/300 (299 tried) — factory 2024-06-17
24.7% Gpt4o 74/300 (289 tried) — moatless 2024-06-17
21.7% Gpt4o 65/300 (290 tried) — appmap-navie 2024-06-15
27.3% Gpt4o 82/300 (102 tried) — MASAI 2024-06-12
26.7% Research_Agent101 80/300 (293 tried) — IBM 2024-06-12
23.7% Starship_gpt4 71/300 (299 tried) — opencsg 2024-05-24
18.0% Gpt4 54/300 (284 tried) — sweagent 2024-04-02
11.7% Claude3opus 35/300 (271 tried) — sweagent 2024-04-02
4.3% Claude3opus 13/300 — rag 2024-04-02
2.7% Gpt4 8/300 — rag 2024-04-02
3.0% Claude2 9/300 (299 tried) — rag 2023-10-10
1.3% Swellama7b 4/300 (292 tried) — rag 2023-10-10
1.0% Swellama13b 3/300 (287 tried) — rag 2023-10-10
0.3% Gpt35 1/300 — rag 2023-10-10

SWE-bench Multilingual 14 submissions

ResolvedModelInstancesCost / taskHarnessRun
67.0% Gemini 3.5 Flash 201/300 (264 tried) $0.63 mini-v2.4.6 2026-09-02
66.3% GPT 5.2 Codex 199/300 $0.66 mini-v2.0.0 2026-02-20
67.7% Minimax 2.5 203/300 (297 tried) $0.10 mini-v2.0.0a0 2026-02-16
72.7% Gemini 3 Flash 218/300 $0.35 mini-v2.0.0a0 2026-02-13
72.0% Claude 4.6 Opus 216/300 $0.66 mini-v2.0.0a0 2026-02-13
70.7% Claude 4.5 Opus 212/300 $0.83 mini-v2.0.0a0 2026-02-13
69.7% GLM 5 209/300 $0.64 mini-v2.0.0a0 2026-02-13
68.7% Gemini 3 Pro 206/300 $1.02 mini-v2.0.0a0 2026-02-13
67.3% Kimi K2 5 202/300 $0.69 mini-v2.0.0a0 2026-02-13
67.0% Claude 4.5 Sonnet 201/300 $0.67 mini-v2.0.0a0 2026-02-13
66.7% GPT 5.2 200/300 $0.54 mini-v2.0.0a0 2026-02-13
64.7% Claude 4.5 Haiku 194/300 $0.38 mini-v2.0.0a0 2026-02-13
59.0% DeepSeek 3.2 177/300 $0.38 mini-v2.0.0a0 2026-02-13
39.7% GPT 5 Mini 119/300 $0.052 mini-v2.0.0a0 2026-02-13

SWE-bench Verified 119 submissions

ResolvedModelInstancesCost / taskHarnessRun
71.8% Gemini 3.5 Flash 359/500 (441 tried) — mini-v2.4.2 2026-09-01
76.8% Claude 4.5 Opus 384/500 $0.75 mini-v2.0.0 2026-02-17
75.8% Gemini 3 Flash 379/500 $0.36 mini-v2.0.0 2026-02-17
75.8% Minimax 2.5 379/500 $0.073 mini-v2.0.0 2026-02-17
75.6% Claude 4.6 Opus 378/500 $0.55 mini-v2.0.0 2026-02-17
72.8% GLM 5 364/500 $0.53 mini-v2.0.0 2026-02-17
72.8% GPT 5.2 364/500 $0.47 mini-v2.0.0 2026-02-17
71.4% Claude 4.5 Sonnet 357/500 $0.66 mini-v2.0.0 2026-02-17
70.8% Kimi K2 5 354/500 $0.15 mini-v2.0.0 2026-02-17
70.0% DeepSeek 3.2 350/500 $0.45 mini-v2.0.0 2026-02-17
66.6% Claude 4.5 Haiku 333/500 $0.33 mini-v2.0.0 2026-02-17
56.2% GPT 5 Mini 281/500 $0.047 mini-v2.0.0 2026-02-17
79.2% Claude Opus 4.5 396/500 (401 tried) — livesweagent 2025-12-15
71.8% GPT 5.2 2025 12 11 359/500 $0.52 mini-v1.17.2 2025-12-11
69.0% GPT 5.2 2025 12 11 345/500 $0.27 mini-v1.17.2 2025-12-11
63.4% Kimi K2 317/500 $0.44 mini-v1.17.2 2025-12-10
56.4% Devstral Small 2512 282/500 $0.24 mini-v1.17.2 2025-12-09
53.8% Devstral 2512 269/500 $0.68 mini-v1.17.2 2025-12-09
79.2% Claude Opus 4.5 396/500 (396 tried) — sonar-foundation-agent 2025-12-05
60.0% DeepSeek V3.2 Reasoner 300/500 $0.028 mini-v1.17.1 2025-12-01
55.4% GLM 4.6 277/500 $0.097 mini-v1.17.1 2025-12-01
77.6% Claude Opus 4.5 388/500 (390 tried) — openhands 2025-11-27
74.4% Claude Opus 4.5 20251101 372/500 $0.72 mini-v1.16.0 2025-11-24
66.0% GPT 5.1 Codex 330/500 $0.59 mini-v1.16.0 2025-11-24
61.0% Minimax M2 305/500 $0.43 mini-v1.17.0 2025-11-24
77.4% Gemini 3 Pro 387/500 (390 tried) — livesweagent 2025-11-20
66.0% GPT 5.1 2025 11 13 330/500 $0.31 mini-v1.15.0 2025-11-20
74.2% Gemini 3 Pro Preview 20251118 371/500 $0.46 mini-v1.15.0 2025-11-18
74.8% Claude Sonnet 4.5 374/500 (377 tried) — sonar-foundation-agent 2025-11-03
73.8% SAGE_OpenHands 369/500 (372 tried) — SalesforceAIResearch 2025-11-03
73.0% SAGE_bash_only 365/500 (366 tried) — SalesforceAIResearch 2025-10-21
74.4% V1.2.1_gpt5 372/500 (373 tried) — Prometheus 2025-10-15
71.2% Kimi_k2 356/500 (357 tried) — Lingxi 2025-10-14
68.2% Glm4 6 341/500 (344 tried) — zai 2025-09-30
71.2% V1.2_gpt5 356/500 (357 tried) — Prometheus 2025-09-29
70.6% Sonnet 4.5 20250929 353/500 $0.56 mini-v1.13.3 2025-09-29
78.8% Doubao_seed_code 394/500 (396 tried) — trae 2025-09-28
57.0% Agent_v2 285/500 (297 tried) — artemis 2025-09-24
60.4% R2E_QwenCoder30BA3B_tts 302/500 (307 tried) — entroPO 2025-09-01
52.2% R2E_QwenCoder30BA3B 261/500 (273 tried) — entroPO 2025-09-01
54.2% GLM 4.5 271/500 $0.30 mini-v1.9.1 2025-08-22
71.8% Gpt5 359/500 (360 tried) — openhands 2025-08-07
65.0% GPT 5 325/500 $0.28 mini-v1.7.0 2025-08-07
59.8% GPT 5 Mini 299/500 $0.035 mini-v1.7.0 2025-08-07
43.8% Kimi K2 Instruct 219/500 $0.53 mini-v1.7.0 2025-08-07
34.8% GPT 5 Nano 174/500 $0.038 mini-v1.7.0 2025-08-07
26.0% GPT Oss 120b 130/500 $0.057 mini-v1.7.0 2025-08-07
42.0% DeepSeek V3 210/500 (242 tried) — SWE-Exp 2025-08-06
53.4% Sweagent_kimi_k2_instruct 267/500 (286 tried) — codesweep 2025-08-04
9.0% Qwen2 5 Coder 32b Instruct 45/500 $0.068 mini-v1.0.0 2025-08-03
67.6% Claude 4 Opus 20250514 338/500 $1.13 mini-v1.0.0 2025-08-02
55.4% Qwen3 Coder 480b A35b Instruct 277/500 $0.25 mini-v1.0.0 2025-08-02
74.8% Ai 374/500 (374 tried) — harness 2025-07-31
64.2% Glm4 5 321/500 (322 tried) — zai 2025-07-28
64.8% Claude Sonnet 4 20250514 324/500 $0.37 mini-v1.0.0 2025-07-26
58.4% O3 2025 04 16 292/500 $0.33 mini-v1.0.0 2025-07-26
53.6% Gemini 2.5 Pro 268/500 $0.29 mini-v1.0.0 2025-07-26
45.0% O4 Mini 2025 04 16 225/500 $0.21 mini-v1.0.0 2025-07-26
38.0% Devstral_small_2507 190/500 (196 tried) — sweagent 2025-07-25
74.6% Claude 4 Sonnet 20250514 373/500 (374 tried) — Lingxi-v1.5 2025-07-20
65.4% Kimi_k2 327/500 (328 tried) — openhands 2025-07-16
71.2% Command 356/500 (356 tried) — qodo 2025-07-15
58.8% R2eagent_tts 294/500 (297 tried) — deepswerl 2025-06-29
42.2% R2eagent 211/500 (256 tried) — deepswerl 2025-06-29
23.2% MCTS Refine 7B 116/500 (155 tried) — agentless 2025-06-27
47.0% Bo8 235/500 (236 tried) — Skywork-SWE-32B+TTS 2025-06-16
70.8% Claude 4 Sonnet 20250514 354/500 (357 tried) — moatless 2025-06-11
70.4% Agent_v1 352/500 (352 tried) — augment 2025-06-10
74.4% Agent_claude 4 Sonnet 372/500 (373 tried) — Refact 2025-06-03
46.0% Co PatcheR 230/500 (230 tried) — patchpilot 2025-05-28
70.4% Claude_4_sonnet 352/500 (353 tried) — openhands 2025-05-24
73.2% Claude 4 Opus 366/500 (366 tried) — tools 2025-05-22
72.4% Claude 4 Sonnet 362/500 (362 tried) — tools 2025-05-22
66.6% Claude 4 Sonnet 20250514 333/500 (333 tried) — sweagent 2025-05-22
46.8% Devstral_small 234/500 (242 tried) — openhands 2025-05-20
68.2% O3 341/500 (462 tried) — cortexa 2025-05-16
70.4% Agent 352/500 (353 tried) — Refact 2025-05-15
66.4% Coder 332/500 (332 tried) — aime 2025-05-14
40.2% Lm_32b 201/500 (205 tried) — sweagent 2025-05-11
70.0% Ai 350/500 (351 tried) — zencoder 2025-04-30
56.6% Claude37 283/500 (283 tried) — swe-rizzo 2025-04-05
65.4% Agent_v0 327/500 (327 tried) — augment 2025-03-16
32.8% Qwen2.5 7b Retriever_Qwen2.5 72b Editor 164/500 (183 tried) — SWE-Fixer 2025-03-06
41.2% Llama3_70b 206/500 (207 tried) — swerl 2025-02-26
62.4% Claude 3.7 Sonnet 312/500 (314 tried) — sweagent 2025-02-25
63.2% Claude 3.7 Sonnet 316/500 (316 tried) — tools 2025-02-24
42.4% Lite_o3_mini 212/500 (222 tried) — agentless 2025-02-14
60.8% 4x_scaled 304/500 (306 tried) — openhands 2025-02-03
44.2% Gemini_2.0_flash_experimental 221/500 (235 tried) — codeshellagent 2025-01-18
64.6% Programmer_o1_crosscheck5 323/500 (324 tried) — wandb 2025-01-17
62.8% Agent_v1.1 314/500 (334 tried) — blackboxai 2025-01-10
60.2% By_interact_claude3.5 301/500 (391 tried) — learn 2025-01-10
62.2% Midwit_claude 3.5 Sonnet_swe Search 311/500 (356 tried) — codestory 2024-12-21
52.2% Jules_gemini_2.0_flash_experimental 261/500 (261 tried) — google 2024-12-12
50.8% Claude 3.5 Sonnet 20241022 254/500 (257 tried) — agentless-1.5 2024-12-02
30.2% Qwen2.5 7b Retriever_Qwen2.5 72b Editor_20241128 151/500 (153 tried) — SWE-Fixer 2024-11-28
32.0% Agent 160/500 (171 tried) — artemis 2024-11-20
38.8% Gpt4o 194/500 (198 tried) — agentless-1.5 2024-10-28
48.6% Swekit 243/500 (244 tried) — composio 2024-10-25
49.0% Claude 3.5 Sonnet Updated 245/500 (262 tried) — tools 2024-10-22
40.6% Claude 3.5 Haiku 203/500 (222 tried) — tools 2024-10-22
40.6% Swekit 203/500 (203 tried) — composio 2024-10-16
28.8% Lingma Swe GPT 72b 144/500 (153 tried) — lingma-agent 2024-10-02
18.2% Lingma Swe GPT 7b 91/500 (179 tried) — lingma-agent 2024-10-02
25.0% Lingma Swe GPT 72b 125/500 (142 tried) — lingma-agent 2024-09-18
10.2% Lingma Swe GPT 7b 51/500 (92 tried) — lingma-agent 2024-09-18
23.2% Gpt4o 116/500 (169 tried) — sweagent 2024-07-28
33.6% Claude3.5sonnet 168/500 (182 tried) — sweagent 2024-06-20
37.0% Code_droid 185/500 — factory 2024-06-17
26.2% Gpt4o 131/500 (494 tried) — appmap-navie 2024-06-15
32.6% Gpt4o 163/500 (189 tried) — MASAI 2024-06-12
22.4% Gpt4 112/500 (472 tried) — sweagent 2024-04-02
15.8% Claude3opus 79/500 (455 tried) — sweagent 2024-04-02
7.0% Claude3opus 35/500 — rag 2024-04-02
2.8% Gpt4 14/500 — rag 2024-04-02
4.4% Claude2 22/500 (499 tried) — rag 2023-10-10
1.4% Swellama7b 7/500 (494 tried) — rag 2023-10-10
1.2% Swellama13b 6/500 (479 tried) — rag 2023-10-10
0.4% Gpt35 2/500 — rag 2023-10-10

Every percentage is instances resolved divided by the full size of the split — 500, 300 instances across 3 scored suites. A submission that chose to run fewer instances is still scored against the whole split, and the figure in brackets is how many it actually attempted.

SWE-bench Multimodal is not shown. The multimodal suite changed size mid-collection: 510 instances in the 2025 release (princeton-nlp/SWE-bench_Multimodal card, Jan 2025) and 480 in the current canonical one (SWE-bench/SWE-bench_Multimodal card, Aug 2026). Every multimodal submission in the experiments repo is a partial 2025-era run of 133-195 instances, so no single denominator is correct both for when the run happened and for the leaderboard a reader compares against today. Excluded rather than scored against a size that no longer applies. (9 submissions withheld — see proof-benchmarks.json.)

1 submissions held back rather than shown: Gemini 3 Pro. Each was withheld for an implausible result — typically a 0% or 100% rate that indicates a broken harness rather than a model score. See proof-benchmarks.json for the reason attached to each.

Rows on different harness versions (mini-v2.0.0 vs mini-v2.4.2) are not directly comparable — the task set and tooling changed. Compare within a version. Cost per task is the total reported spend divided by instances attempted.

The models that have no public benchmark

Benchmarks lag launches. There are 2,347 models currently shipping, of which 1,412 were released in 2026. Only 174 appear in any public benchmark suite we can read. The rest are unmeasured.

That gap is the honest state of the field, not a CacheSphere omission. If you are weighing a model that launched this month, nobody has published independent numbers for it yet — treat vendor claims as claims.

Full list with published price, context window and release date: /api/model-landscape.json (238 KB, 2,347 rows). Source: models.dev.

Methodology & raw data

Where these numbers come from

Taken from SWE-bench/experiments, path evaluation/verified. Where a submission ships only per_instance_details.json, the resolve rate is computed as instances marked resolved divided by instances attempted. Nothing is estimated, extrapolated, or smoothed.

What we refuse to show

  • Submissions whose reported result is implausible are withheld with a stated reason, not quietly dropped.
  • We do not run our own benchmark to fill a gap, and we do not present a vendor claim as a measured result.
  • Rows are never merged across harness versions to manufacture a ranking.
  • Read how benchmark evidence is collected and reviewed before citing numbers externally.

Machine-readable artifacts