Rangering fra Arena.ai, blindtester og agentoppgaver.
Agentiske oppgaver med flerstegsresonnering og verktøybruk.
| # | Modell | Leverandør | Netto forbedring |
|---|---|---|---|
| 01 | Claude Fable 5 (High) | Anthropic | +14,07 % |
| 02 | Claude Opus 4.8 (Thinking) | Anthropic | +9,90 % |
| 03 | GPT 5.5 (xHigh) | OpenAI | +8,39 % |
| 04 | Claude Opus 4.7 | Anthropic | +8,16 % |
| 05 | Claude Opus 4.7 (Thinking) | Anthropic | +8,07 % |
| 06 | Claude Sonnet 5 (Thinking) | Anthropic | +7,38 % |
| 07 | GPT 5.5 (High) | OpenAI | +7,16 % |
| 08 | GPT 5.4 (High) | OpenAI | +6,73 % |
| 09 | GLM 5.2 (Max) | Z.ai | +6,72 % |
| 10 | GPT 5.5 | OpenAI | +6,55 % |
| 11 | Claude Opus 4.6 | Anthropic | +6,23 % |
| 12 | Claude Opus 4.8 | Anthropic | +4,69 % |
| 13 | Claude Sonnet 4.6 | Anthropic | +2,53 % |
| 14 | GLM 5.1 | Z.ai | +1,67 % |
| 15 | kimi-k2.7-code | Moonshot | +0,74 % |
| 16 | Gemini 3.1 Pro Preview | −0,55 % | |
| 17 | DeepSeek V4 Flash | DeepSeek | −1,00 % |
| 18 | Kimi K2.7 Code | Moonshot | −1,22 % |
| 19 | kimi-k2.6 | Moonshot | −1,62 % |
| 20 | Kimi K2.6 | Moonshot | −1,84 % |
| 21 | Minimax M3 | MiniMax | −2,45 % |
| 22 | DeepSeek V4 Pro | DeepSeek | −2,67 % |
| 23 | minimax-m3 | MiniMax | −4,48 % |
| 24 | Qwen 3.6 Plus | Alibaba | −4,77 % |
| 25 | gemini-3.5-flash | −6,58 % |