Evidence, not our own benchmarks
Every number below was run by somebody else and published openly. We did not run any of it. What CacheSphere adds is the comparison.
Sources: the SWE-bench project's own experiments repository. Each model links
to the exact submission directory its figures come from, so you can audit any row.
What the data says
On the same harness (mini-v2.0.0), Claude 4.5 Opus resolved the most real issues — 384 of the 500-instance suite, or 76.8%. Minimax 2.5 came within 1.0 points of that at $0.073 per task against $0.75 — roughly 10× cheaper.
Two columns matter. Resolve rate is always measured against the whole split, so a submission that attempted fewer instances still counts every instance it skipped. Cost per task is the spend the submitting harness reported, divided by instances attempted. The expensive model wins on rate; the cheap one wins on almost everything else.
Every submission we found
256 submissions across 179 distinct models, 4 benchmark suites, from 2 independent publishers. Every row links to the exact submission it came from.
Aider polyglot 64 submissions
| Resolved | Model | Instances | Cost / task | Harness | Run |
|---|---|---|---|---|---|
| 88.0% | gpt-5 | 225 | — | diff | 2025-08-23 |
| 84.9% | o3-pro | 225 | — | diff | 2025-06-28 |
| 83.1% | gemini-2.5-pro-preview-06-05 (32k think) | 225 | — | diff-fenced | 2025-06-06 |
| 81.3% | o3 | 225 | — | diff | 2025-06-25 |
| 79.6% | grok-4 | 225 | — | diff | 2025-07-11 |
| 79.1% | gemini-2.5-pro-preview-06-05 (default think) | 225 | — | diff-fenced | 2025-06-06 |
| 78.2% | o3 (high) + gpt-4.1 | 225 | — | architect | 2025-06-27 |
| 76.9% | Gemini 2.5 Pro Preview 05-06 | 225 | — | diff-fenced | 2025-05-07 |
| 74.2% | DeepSeek-V3.2-Exp (Reasoner) | 225 | — | diff | 2025-10-03 |
| 72.9% | Gemini 2.5 Pro Preview 03-25 | 225 | — | diff-fenced | 2025-04-12 |
| 72.0% | o4-mini | 225 | — | diff | 2025-04-16 |
| 72.0% | claude-opus-4-20250514 (32k thinking) | 225 | — | diff | 2025-05-25 |
| 71.4% | DeepSeek R1 (0528) | 225 | — | diff | 2025-06-06 |
| 70.7% | claude-opus-4-20250514 (no think) | 225 | — | diff | 2025-05-25 |
| 70.2% | DeepSeek-V3.2-Exp (Chat) | 225 | — | diff | 2025-10-03 |
| 64.9% | claude-3-7-sonnet-20250219 (32k thinking tokens) | 225 | — | diff | 2025-02-24 |
| 64.0% | DeepSeek R1 + claude-3-5-sonnet-20241022 | 225 | — | architect | 2025-01-23 |
| 61.7% | o1-2024-12-17 | 225 | — | diff | 2024-12-21 |
| 61.3% | claude-sonnet-4-20250514 (32k thinking) | 225 | — | diff | 2025-05-24 |
| 60.4% | o3-mini | 225 | — | diff | 2025-01-31 |
| 60.4% | claude-3-7-sonnet-20250219 (no thinking) | 225 | — | diff | 2025-02-24 |
| 59.6% | Qwen3 235B A22B diff, no think, Alibaba API | 225 | — | diff | 2025-05-09 |
| 59.1% | Kimi K2 | 225 | — | diff | 2025-07-17 |
| 56.9% | DeepSeek R1 | 225 | — | diff | 2025-01-20 |
| 56.4% | claude-sonnet-4-20250514 (no thinking) | 225 | — | diff | 2025-05-24 |
| 55.1% | DeepSeek V3 (0324) | 225 | — | diff | 2025-03-24 |
| 55.1% | gemini-2.5-flash-preview-05-20 (24k think) | 225 | — | diff | 2025-05-25 |
| 54.7% | Quasar Alpha | 225 | — | diff | 2025-04-04 |
| 53.3% | Grok 3 Beta | 225 | — | diff | 2025-04-10 |
| 52.9% | Optimus Alpha | 225 | — | diff | 2025-04-10 |
| 52.4% | gpt-4.1 | 225 | — | diff | 2025-04-14 |
| 51.6% | claude-3-5-sonnet-20241022 | 225 | — | diff | 2025-01-17 |
| 49.3% | Grok 3 Mini Beta | 225 | — | whole | 2025-04-10 |
| 48.4% | DeepSeek Chat V3 (prev) | 225 | — | diff | 2024-12-25 |
| 47.1% | gemini-2.5-flash-preview-04-17 | 225 | — | diff | 2025-04-20 |
| 45.3% | chatgpt-4o-latest (2025-03-29) | 225 | — | diff | 2025-03-29 |
| 44.9% | gpt-4.5-preview | 225 | — | diff | 2025-02-27 |
| 44.0% | gemini-2.5-flash-preview-05-20 (no think) | 225 | — | diff | 2025-05-26 |
| 41.8% | gpt-oss-120b | 225 | — | diff | 2025-08-06 |
| 40.0% | Qwen3 32B | 225 | — | diff | 2025-05-08 |
| 38.2% | gemini-exp-1206 | 225 | — | whole | 2024-12-22 |
| 35.6% | Gemini 2.0 Pro exp-02-05 | 225 | — | whole | 2025-02-25 |
| 32.9% | o1-mini-2024-09-12 | 225 | — | whole | 2024-12-22 |
| 32.4% | gpt-4.1-mini | 225 | — | diff | 2025-04-14 |
| 28.0% | claude-3-5-haiku-20241022 | 225 | — | diff | 2024-12-21 |
| 27.1% | chatgpt-4o-latest (2025-02-15) | 225 | — | diff | 2025-02-15 |
| 26.2% | QwQ-32B + Qwen 2.5 Coder Instruct | 225 | — | architect | 2025-03-07 |
| 23.1% | gpt-4o-2024-08-06 | 225 | — | diff | 2024-12-30 |
| 22.2% | gemini-2.0-flash-exp | 225 | — | whole | 2024-12-22 |
| 21.8% | qwen-max-2025-01-25 | 225 | — | diff | 2025-01-28 |
| 20.9% | QwQ-32B | 225 | — | diff | 2025-03-06 |
| 18.2% | gpt-4o-2024-11-20 | 225 | — | diff | 2024-12-30 |
| 18.2% | gemini-2.0-flash-thinking-exp-01-21 | 225 | — | diff | 2025-01-21 |
| 17.8% | DeepSeek Chat V2.5 | 225 | — | diff | 2024-12-21 |
| 16.4% | Qwen2.5-Coder-32B-Instruct | 225 | — | whole | 2024-12-26 |
| 15.6% | Llama 4 Maverick | 225 | — | whole | 2025-04-06 |
| 12.9% | yi-lightning | 225 | — | whole | 2024-12-23 |
| 12.0% | command-a-03-2025-quality | 225 | — | whole | 2025-03-14 |
| 11.1% | Codestral 25.01 | 225 | — | whole | 2025-01-13 |
| 10.2% | openhands-lm-32b-v0.1 | 225 | — | whole | 2025-04-19 |
| 8.9% | gpt-4.1-nano | 225 | — | whole | 2025-04-14 |
| 8.0% | Qwen2.5-Coder-32B-Instruct | 225 | — | diff | 2024-12-22 |
| 4.9% | gemma-3-27b-it | 225 | — | whole | 2025-03-15 |
| 3.6% | gpt-4o-mini-2024-07-18 | 225 | — | whole | 2024-12-21 |
SWE-bench Lite 59 submissions
| Resolved | Model | Instances | Cost / task | Harness | Run |
|---|---|---|---|---|---|
| 51.3% | Claude 3.5 Sonnet 20241022 | 154/300 (154 tried) | — | isea | 2025-09-11 |
| 58.3% | Claude 4 Sonnet 20250514 | 175/300 (179 tried) | — | KGCompass | 2025-09-06 |
| 49.7% | R2E_QwenCoder30BA3B_tts | 149/300 (149 tried) | — | entroPO | 2025-09-01 |
| 45.0% | R2E_QwenCoder30BA3B | 135/300 (137 tried) | — | entroPO | 2025-09-01 |
| 16.3% | MCTS Refine 7B | 49/300 (87 tried) | — | agentless | 2025-06-27 |
| 60.3% | Claude 4 Sonnet 20250514 | 181/300 (181 tried) | — | ExpeRepair-v1 | 2025-06-25 |
| 51.7% | Multi V1_Claude3.7Sonnet_Gemini2.5Pro | 155/300 (155 tried) | — | SemAgent | 2025-06-25 |
| 46.0% | Claude 3.5 Sonnet 20241022 | 138/300 (147 tried) | — | KGCompass | 2025-06-19 |
| 36.7% | DeepSeek V3 | 110/300 (116 tried) | — | KGCompass | 2025-06-09 |
| 56.7% | Claude 4 Sonnet 20250514 | 170/300 (171 tried) | — | sweagent | 2025-05-26 |
| 42.7% | Claude 3.5 Sonnet 20241022 | 128/300 (129 tried) | — | Lingxi | 2025-05-09 |
| 60.0% | Agent | 180/300 (181 tried) | — | Refact | 2025-04-25 |
| 24.7% | Qwen2.5 7b Retriever_Qwen2.5 72b Editor | 74/300 (78 tried) | — | SWE-Fixer | 2025-03-06 |
| 48.0% | Claude 3.7 Sonnet 20250219 | 144/300 (146 tried) | — | sweagent | 2025-02-26 |
| 32.3% | Lite_o3_mini | 97/300 (103 tried) | — | agentless | 2025-02-14 |
| 30.3% | O3mini | 91/300 (130 tried) | — | aegis | 2025-02-07 |
| 47.0% | Agent_claude_3.5_sonnet_deepseek_r1 | 141/300 (141 tried) | — | dars | 2025-02-05 |
| 39.0% | Claude 3.5 Sonnet 20241022 | 117/300 (128 tried) | — | moatless | 2025-01-14 |
| 39.7% | Gpt4o | 119/300 (133 tried) | — | OpenCSG-Starship-Agentic-Coder | 2025-01-13 |
| 30.7% | Deepseek_v3 | 92/300 (105 tried) | — | moatless | 2025-01-11 |
| 37.0% | Codes_claude 3.5 Sonnet 20241022 | 111/300 (224 tried) | — | patched | 2025-01-04 |
| 49.0% | Agent_v1 | 147/300 (159 tried) | — | blackboxai | 2024-12-20 |
| 41.3% | Claude 3.5 Sonnet 20241022 | 124/300 (129 tried) | — | PatchKitty-0.9 | 2024-12-20 |
| 44.7% | Sonnet_v1 | 134/300 (208 tried) | — | kodu | 2024-12-07 |
| 40.7% | Claude 3.5 Sonnet 20241022 | 122/300 (124 tried) | — | agentless-1.5 | 2024-12-02 |
| 23.3% | Qwen2.5 7b Retriever_Qwen2.5 72b Editor_20241128 | 70/300 (70 tried) | — | SWE-Fixer | 2024-11-28 |
| 48.3% | Codefixer_agent | 145/300 (160 tried) | — | globant | 2024-11-27 |
| 38.3% | Claude 3.5 Sonnet 20241022 | 115/300 (126 tried) | — | moatless | 2024-11-17 |
| 28.0% | Gpt4o | 84/300 (91 tried) | — | reproducedRG | 2024-11-17 |
| 31.3% | Gpt4o | 94/300 (106 tried) | — | codeshelltester | 2024-11-11 |
| 41.0% | Swekit | 123/300 (124 tried) | — | composio | 2024-10-30 |
| 32.0% | Gpt4o | 96/300 (100 tried) | — | agentless-1.5 | 2024-10-28 |
| 25.3% | Lite1 | 76/300 (119 tried) | — | hyperagent | 2024-09-25 |
| 30.0% | Gpt4o | 90/300 (117 tried) | — | infant | 2024-09-08 |
| 21.7% | Mixed | 65/300 (96 tried) | — | autose | 2024-08-28 |
| 29.7% | Gpt4o | 89/300 (96 tried) | — | RepoGraph | 2024-08-08 |
| 18.3% | Gpt4o | 55/300 (94 tried) | — | sweagent | 2024-07-28 |
| 26.7% | Codeact_v1.8_claude35sonnet | 80/300 (113 tried) | — | opendevin | 2024-07-25 |
| 27.7% | Gpt4o | 83/300 | — | sima | 2024-07-06 |
| 43.0% | Aide_mixed | 129/300 (273 tried) | — | codestory | 2024-07-02 |
| 27.3% | Gpt4o | 82/300 (292 tried) | — | agentless | 2024-06-30 |
| 38.0% | Mentatbot_gpt4o | 114/300 (296 tried) | — | abanteai | 2024-06-27 |
| 26.7% | Claude35sonnet | 80/300 (294 tried) | — | moatless | 2024-06-23 |
| 33.0% | Agent | 99/300 (292 tried) | — | Lingma | 2024-06-22 |
| 23.0% | Claude3.5sonnet | 69/300 (81 tried) | — | sweagent | 2024-06-20 |
| 31.3% | Code_droid | 94/300 (299 tried) | — | factory | 2024-06-17 |
| 24.7% | Gpt4o | 74/300 (289 tried) | — | moatless | 2024-06-17 |
| 21.7% | Gpt4o | 65/300 (290 tried) | — | appmap-navie | 2024-06-15 |
| 27.3% | Gpt4o | 82/300 (102 tried) | — | MASAI | 2024-06-12 |
| 26.7% | Research_Agent101 | 80/300 (293 tried) | — | IBM | 2024-06-12 |
| 23.7% | Starship_gpt4 | 71/300 (299 tried) | — | opencsg | 2024-05-24 |
| 18.0% | Gpt4 | 54/300 (284 tried) | — | sweagent | 2024-04-02 |
| 11.7% | Claude3opus | 35/300 (271 tried) | — | sweagent | 2024-04-02 |
| 4.3% | Claude3opus | 13/300 | — | rag | 2024-04-02 |
| 2.7% | Gpt4 | 8/300 | — | rag | 2024-04-02 |
| 3.0% | Claude2 | 9/300 (299 tried) | — | rag | 2023-10-10 |
| 1.3% | Swellama7b | 4/300 (292 tried) | — | rag | 2023-10-10 |
| 1.0% | Swellama13b | 3/300 (287 tried) | — | rag | 2023-10-10 |
| 0.3% | Gpt35 | 1/300 | — | rag | 2023-10-10 |
SWE-bench Multilingual 14 submissions
| Resolved | Model | Instances | Cost / task | Harness | Run |
|---|---|---|---|---|---|
| 67.0% | Gemini 3.5 Flash | 201/300 (264 tried) | $0.63 | mini-v2.4.6 | 2026-09-02 |
| 66.3% | GPT 5.2 Codex | 199/300 | $0.66 | mini-v2.0.0 | 2026-02-20 |
| 67.7% | Minimax 2.5 | 203/300 (297 tried) | $0.10 | mini-v2.0.0a0 | 2026-02-16 |
| 72.7% | Gemini 3 Flash | 218/300 | $0.35 | mini-v2.0.0a0 | 2026-02-13 |
| 72.0% | Claude 4.6 Opus | 216/300 | $0.66 | mini-v2.0.0a0 | 2026-02-13 |
| 70.7% | Claude 4.5 Opus | 212/300 | $0.83 | mini-v2.0.0a0 | 2026-02-13 |
| 69.7% | GLM 5 | 209/300 | $0.64 | mini-v2.0.0a0 | 2026-02-13 |
| 68.7% | Gemini 3 Pro | 206/300 | $1.02 | mini-v2.0.0a0 | 2026-02-13 |
| 67.3% | Kimi K2 5 | 202/300 | $0.69 | mini-v2.0.0a0 | 2026-02-13 |
| 67.0% | Claude 4.5 Sonnet | 201/300 | $0.67 | mini-v2.0.0a0 | 2026-02-13 |
| 66.7% | GPT 5.2 | 200/300 | $0.54 | mini-v2.0.0a0 | 2026-02-13 |
| 64.7% | Claude 4.5 Haiku | 194/300 | $0.38 | mini-v2.0.0a0 | 2026-02-13 |
| 59.0% | DeepSeek 3.2 | 177/300 | $0.38 | mini-v2.0.0a0 | 2026-02-13 |
| 39.7% | GPT 5 Mini | 119/300 | $0.052 | mini-v2.0.0a0 | 2026-02-13 |
SWE-bench Verified 119 submissions
| Resolved | Model | Instances | Cost / task | Harness | Run |
|---|---|---|---|---|---|
| 71.8% | Gemini 3.5 Flash | 359/500 (441 tried) | — | mini-v2.4.2 | 2026-09-01 |
| 76.8% | Claude 4.5 Opus | 384/500 | $0.75 | mini-v2.0.0 | 2026-02-17 |
| 75.8% | Gemini 3 Flash | 379/500 | $0.36 | mini-v2.0.0 | 2026-02-17 |
| 75.8% | Minimax 2.5 | 379/500 | $0.073 | mini-v2.0.0 | 2026-02-17 |
| 75.6% | Claude 4.6 Opus | 378/500 | $0.55 | mini-v2.0.0 | 2026-02-17 |
| 72.8% | GLM 5 | 364/500 | $0.53 | mini-v2.0.0 | 2026-02-17 |
| 72.8% | GPT 5.2 | 364/500 | $0.47 | mini-v2.0.0 | 2026-02-17 |
| 71.4% | Claude 4.5 Sonnet | 357/500 | $0.66 | mini-v2.0.0 | 2026-02-17 |
| 70.8% | Kimi K2 5 | 354/500 | $0.15 | mini-v2.0.0 | 2026-02-17 |
| 70.0% | DeepSeek 3.2 | 350/500 | $0.45 | mini-v2.0.0 | 2026-02-17 |
| 66.6% | Claude 4.5 Haiku | 333/500 | $0.33 | mini-v2.0.0 | 2026-02-17 |
| 56.2% | GPT 5 Mini | 281/500 | $0.047 | mini-v2.0.0 | 2026-02-17 |
| 79.2% | Claude Opus 4.5 | 396/500 (401 tried) | — | livesweagent | 2025-12-15 |
| 71.8% | GPT 5.2 2025 12 11 | 359/500 | $0.52 | mini-v1.17.2 | 2025-12-11 |
| 69.0% | GPT 5.2 2025 12 11 | 345/500 | $0.27 | mini-v1.17.2 | 2025-12-11 |
| 63.4% | Kimi K2 | 317/500 | $0.44 | mini-v1.17.2 | 2025-12-10 |
| 56.4% | Devstral Small 2512 | 282/500 | $0.24 | mini-v1.17.2 | 2025-12-09 |
| 53.8% | Devstral 2512 | 269/500 | $0.68 | mini-v1.17.2 | 2025-12-09 |
| 79.2% | Claude Opus 4.5 | 396/500 (396 tried) | — | sonar-foundation-agent | 2025-12-05 |
| 60.0% | DeepSeek V3.2 Reasoner | 300/500 | $0.028 | mini-v1.17.1 | 2025-12-01 |
| 55.4% | GLM 4.6 | 277/500 | $0.097 | mini-v1.17.1 | 2025-12-01 |
| 77.6% | Claude Opus 4.5 | 388/500 (390 tried) | — | openhands | 2025-11-27 |
| 74.4% | Claude Opus 4.5 20251101 | 372/500 | $0.72 | mini-v1.16.0 | 2025-11-24 |
| 66.0% | GPT 5.1 Codex | 330/500 | $0.59 | mini-v1.16.0 | 2025-11-24 |
| 61.0% | Minimax M2 | 305/500 | $0.43 | mini-v1.17.0 | 2025-11-24 |
| 77.4% | Gemini 3 Pro | 387/500 (390 tried) | — | livesweagent | 2025-11-20 |
| 66.0% | GPT 5.1 2025 11 13 | 330/500 | $0.31 | mini-v1.15.0 | 2025-11-20 |
| 74.2% | Gemini 3 Pro Preview 20251118 | 371/500 | $0.46 | mini-v1.15.0 | 2025-11-18 |
| 74.8% | Claude Sonnet 4.5 | 374/500 (377 tried) | — | sonar-foundation-agent | 2025-11-03 |
| 73.8% | SAGE_OpenHands | 369/500 (372 tried) | — | SalesforceAIResearch | 2025-11-03 |
| 73.0% | SAGE_bash_only | 365/500 (366 tried) | — | SalesforceAIResearch | 2025-10-21 |
| 74.4% | V1.2.1_gpt5 | 372/500 (373 tried) | — | Prometheus | 2025-10-15 |
| 71.2% | Kimi_k2 | 356/500 (357 tried) | — | Lingxi | 2025-10-14 |
| 68.2% | Glm4 6 | 341/500 (344 tried) | — | zai | 2025-09-30 |
| 71.2% | V1.2_gpt5 | 356/500 (357 tried) | — | Prometheus | 2025-09-29 |
| 70.6% | Sonnet 4.5 20250929 | 353/500 | $0.56 | mini-v1.13.3 | 2025-09-29 |
| 78.8% | Doubao_seed_code | 394/500 (396 tried) | — | trae | 2025-09-28 |
| 57.0% | Agent_v2 | 285/500 (297 tried) | — | artemis | 2025-09-24 |
| 60.4% | R2E_QwenCoder30BA3B_tts | 302/500 (307 tried) | — | entroPO | 2025-09-01 |
| 52.2% | R2E_QwenCoder30BA3B | 261/500 (273 tried) | — | entroPO | 2025-09-01 |
| 54.2% | GLM 4.5 | 271/500 | $0.30 | mini-v1.9.1 | 2025-08-22 |
| 71.8% | Gpt5 | 359/500 (360 tried) | — | openhands | 2025-08-07 |
| 65.0% | GPT 5 | 325/500 | $0.28 | mini-v1.7.0 | 2025-08-07 |
| 59.8% | GPT 5 Mini | 299/500 | $0.035 | mini-v1.7.0 | 2025-08-07 |
| 43.8% | Kimi K2 Instruct | 219/500 | $0.53 | mini-v1.7.0 | 2025-08-07 |
| 34.8% | GPT 5 Nano | 174/500 | $0.038 | mini-v1.7.0 | 2025-08-07 |
| 26.0% | GPT Oss 120b | 130/500 | $0.057 | mini-v1.7.0 | 2025-08-07 |
| 42.0% | DeepSeek V3 | 210/500 (242 tried) | — | SWE-Exp | 2025-08-06 |
| 53.4% | Sweagent_kimi_k2_instruct | 267/500 (286 tried) | — | codesweep | 2025-08-04 |
| 9.0% | Qwen2 5 Coder 32b Instruct | 45/500 | $0.068 | mini-v1.0.0 | 2025-08-03 |
| 67.6% | Claude 4 Opus 20250514 | 338/500 | $1.13 | mini-v1.0.0 | 2025-08-02 |
| 55.4% | Qwen3 Coder 480b A35b Instruct | 277/500 | $0.25 | mini-v1.0.0 | 2025-08-02 |
| 74.8% | Ai | 374/500 (374 tried) | — | harness | 2025-07-31 |
| 64.2% | Glm4 5 | 321/500 (322 tried) | — | zai | 2025-07-28 |
| 64.8% | Claude Sonnet 4 20250514 | 324/500 | $0.37 | mini-v1.0.0 | 2025-07-26 |
| 58.4% | O3 2025 04 16 | 292/500 | $0.33 | mini-v1.0.0 | 2025-07-26 |
| 53.6% | Gemini 2.5 Pro | 268/500 | $0.29 | mini-v1.0.0 | 2025-07-26 |
| 45.0% | O4 Mini 2025 04 16 | 225/500 | $0.21 | mini-v1.0.0 | 2025-07-26 |
| 38.0% | Devstral_small_2507 | 190/500 (196 tried) | — | sweagent | 2025-07-25 |
| 74.6% | Claude 4 Sonnet 20250514 | 373/500 (374 tried) | — | Lingxi-v1.5 | 2025-07-20 |
| 65.4% | Kimi_k2 | 327/500 (328 tried) | — | openhands | 2025-07-16 |
| 71.2% | Command | 356/500 (356 tried) | — | qodo | 2025-07-15 |
| 58.8% | R2eagent_tts | 294/500 (297 tried) | — | deepswerl | 2025-06-29 |
| 42.2% | R2eagent | 211/500 (256 tried) | — | deepswerl | 2025-06-29 |
| 23.2% | MCTS Refine 7B | 116/500 (155 tried) | — | agentless | 2025-06-27 |
| 47.0% | Bo8 | 235/500 (236 tried) | — | Skywork-SWE-32B+TTS | 2025-06-16 |
| 70.8% | Claude 4 Sonnet 20250514 | 354/500 (357 tried) | — | moatless | 2025-06-11 |
| 70.4% | Agent_v1 | 352/500 (352 tried) | — | augment | 2025-06-10 |
| 74.4% | Agent_claude 4 Sonnet | 372/500 (373 tried) | — | Refact | 2025-06-03 |
| 46.0% | Co PatcheR | 230/500 (230 tried) | — | patchpilot | 2025-05-28 |
| 70.4% | Claude_4_sonnet | 352/500 (353 tried) | — | openhands | 2025-05-24 |
| 73.2% | Claude 4 Opus | 366/500 (366 tried) | — | tools | 2025-05-22 |
| 72.4% | Claude 4 Sonnet | 362/500 (362 tried) | — | tools | 2025-05-22 |
| 66.6% | Claude 4 Sonnet 20250514 | 333/500 (333 tried) | — | sweagent | 2025-05-22 |
| 46.8% | Devstral_small | 234/500 (242 tried) | — | openhands | 2025-05-20 |
| 68.2% | O3 | 341/500 (462 tried) | — | cortexa | 2025-05-16 |
| 70.4% | Agent | 352/500 (353 tried) | — | Refact | 2025-05-15 |
| 66.4% | Coder | 332/500 (332 tried) | — | aime | 2025-05-14 |
| 40.2% | Lm_32b | 201/500 (205 tried) | — | sweagent | 2025-05-11 |
| 70.0% | Ai | 350/500 (351 tried) | — | zencoder | 2025-04-30 |
| 56.6% | Claude37 | 283/500 (283 tried) | — | swe-rizzo | 2025-04-05 |
| 65.4% | Agent_v0 | 327/500 (327 tried) | — | augment | 2025-03-16 |
| 32.8% | Qwen2.5 7b Retriever_Qwen2.5 72b Editor | 164/500 (183 tried) | — | SWE-Fixer | 2025-03-06 |
| 41.2% | Llama3_70b | 206/500 (207 tried) | — | swerl | 2025-02-26 |
| 62.4% | Claude 3.7 Sonnet | 312/500 (314 tried) | — | sweagent | 2025-02-25 |
| 63.2% | Claude 3.7 Sonnet | 316/500 (316 tried) | — | tools | 2025-02-24 |
| 42.4% | Lite_o3_mini | 212/500 (222 tried) | — | agentless | 2025-02-14 |
| 60.8% | 4x_scaled | 304/500 (306 tried) | — | openhands | 2025-02-03 |
| 44.2% | Gemini_2.0_flash_experimental | 221/500 (235 tried) | — | codeshellagent | 2025-01-18 |
| 64.6% | Programmer_o1_crosscheck5 | 323/500 (324 tried) | — | wandb | 2025-01-17 |
| 62.8% | Agent_v1.1 | 314/500 (334 tried) | — | blackboxai | 2025-01-10 |
| 60.2% | By_interact_claude3.5 | 301/500 (391 tried) | — | learn | 2025-01-10 |
| 62.2% | Midwit_claude 3.5 Sonnet_swe Search | 311/500 (356 tried) | — | codestory | 2024-12-21 |
| 52.2% | Jules_gemini_2.0_flash_experimental | 261/500 (261 tried) | — | 2024-12-12 | |
| 50.8% | Claude 3.5 Sonnet 20241022 | 254/500 (257 tried) | — | agentless-1.5 | 2024-12-02 |
| 30.2% | Qwen2.5 7b Retriever_Qwen2.5 72b Editor_20241128 | 151/500 (153 tried) | — | SWE-Fixer | 2024-11-28 |
| 32.0% | Agent | 160/500 (171 tried) | — | artemis | 2024-11-20 |
| 38.8% | Gpt4o | 194/500 (198 tried) | — | agentless-1.5 | 2024-10-28 |
| 48.6% | Swekit | 243/500 (244 tried) | — | composio | 2024-10-25 |
| 49.0% | Claude 3.5 Sonnet Updated | 245/500 (262 tried) | — | tools | 2024-10-22 |
| 40.6% | Claude 3.5 Haiku | 203/500 (222 tried) | — | tools | 2024-10-22 |
| 40.6% | Swekit | 203/500 (203 tried) | — | composio | 2024-10-16 |
| 28.8% | Lingma Swe GPT 72b | 144/500 (153 tried) | — | lingma-agent | 2024-10-02 |
| 18.2% | Lingma Swe GPT 7b | 91/500 (179 tried) | — | lingma-agent | 2024-10-02 |
| 25.0% | Lingma Swe GPT 72b | 125/500 (142 tried) | — | lingma-agent | 2024-09-18 |
| 10.2% | Lingma Swe GPT 7b | 51/500 (92 tried) | — | lingma-agent | 2024-09-18 |
| 23.2% | Gpt4o | 116/500 (169 tried) | — | sweagent | 2024-07-28 |
| 33.6% | Claude3.5sonnet | 168/500 (182 tried) | — | sweagent | 2024-06-20 |
| 37.0% | Code_droid | 185/500 | — | factory | 2024-06-17 |
| 26.2% | Gpt4o | 131/500 (494 tried) | — | appmap-navie | 2024-06-15 |
| 32.6% | Gpt4o | 163/500 (189 tried) | — | MASAI | 2024-06-12 |
| 22.4% | Gpt4 | 112/500 (472 tried) | — | sweagent | 2024-04-02 |
| 15.8% | Claude3opus | 79/500 (455 tried) | — | sweagent | 2024-04-02 |
| 7.0% | Claude3opus | 35/500 | — | rag | 2024-04-02 |
| 2.8% | Gpt4 | 14/500 | — | rag | 2024-04-02 |
| 4.4% | Claude2 | 22/500 (499 tried) | — | rag | 2023-10-10 |
| 1.4% | Swellama7b | 7/500 (494 tried) | — | rag | 2023-10-10 |
| 1.2% | Swellama13b | 6/500 (479 tried) | — | rag | 2023-10-10 |
| 0.4% | Gpt35 | 2/500 | — | rag | 2023-10-10 |
Every percentage is instances resolved divided by the full size of the split — 500, 300 instances across 3 scored suites. A submission that chose to run fewer instances is still scored against the whole split, and the figure in brackets is how many it actually attempted.
SWE-bench Multimodal is not shown. The multimodal suite changed size mid-collection: 510 instances in the 2025 release (princeton-nlp/SWE-bench_Multimodal card, Jan 2025) and 480 in the current canonical one (SWE-bench/SWE-bench_Multimodal card, Aug 2026). Every multimodal submission in the experiments repo is a partial 2025-era run of 133-195 instances, so no single denominator is correct both for when the run happened and for the leaderboard a reader compares against today. Excluded rather than scored against a size that no longer applies. (9 submissions withheld — see proof-benchmarks.json.)
1 submissions held back rather than shown: Gemini 3 Pro. Each was withheld for an implausible result — typically a 0% or 100% rate that indicates a broken harness rather than a model score. See proof-benchmarks.json for the reason attached to each.
Rows on different harness versions (mini-v2.0.0 vs mini-v2.4.2)
are not directly comparable — the task set and tooling changed. Compare
within a version. Cost per task is the total reported spend divided by instances attempted.
The models that have no public benchmark
Benchmarks lag launches. There are 2,347 models currently shipping, of which 1,412 were released in 2026. Only 174 appear in any public benchmark suite we can read. The rest are unmeasured.
That gap is the honest state of the field, not a CacheSphere omission. If you are weighing a model that launched this month, nobody has published independent numbers for it yet — treat vendor claims as claims.
- 2,347 current models tracked
- 600 sold by more than one provider — same model, different price
- 1,412 released in 2026
- 174 with public benchmark evidence
Full list with published price, context window and release date: /api/model-landscape.json (238 KB, 2,347 rows). Source: models.dev.
Methodology & raw data
Where these numbers come from
Taken from SWE-bench/experiments, path evaluation/verified. Where a submission ships only per_instance_details.json, the resolve rate is computed as instances marked resolved divided by instances attempted. Nothing is estimated, extrapolated, or smoothed.
What we refuse to show
- Submissions whose reported result is implausible are withheld with a stated reason, not quietly dropped.
- We do not run our own benchmark to fill a gap, and we do not present a vendor claim as a measured result.
- Rows are never merged across harness versions to manufacture a ranking.
- Read how benchmark evidence is collected and reviewed before citing numbers externally.