Compare
Compare models and agents
Select two to four subjects and compare intelligence, coding, agent power, speed, cost, context, vision and tool use.
Select 2-4 models or agents
The MVP comparison uses normalized demo scores.
| Metric | GPT-5.5 | Claude Opus / Sonnet | Codex | Claude Code |
|---|---|---|---|---|
| Provider | OpenAI | Anthropic | OpenAI | Anthropic |
| Intelligence | 96 | 95 | Agent dependent | Agent dependent |
| Coding | 95 | 93 | 94 | 92 |
| Agent Power | 94 | 91 | 92.3 | 90.5 |
| Speed | 78 | 82 | Host dependent | Host dependent |
| Cost | 72 / Premium | 76 / High | 78 / Medium | 74 / High |
| Context Window | 128K tokens | 200K tokens | Model dependent | Model dependent |
| Vision | 92 | 88 | Model dependent | Model dependent |
| Tool Use | Yes | Yes | 96 | 90 |
| Strengths | Strong tool use, Reliable reasoning, Broad workflow coverage | Long-context synthesis, Code review, Document reasoning | Repo edits, Debugging loops, Build/test driven coding | Careful refactors, Long reviews, Spec-heavy coding |
| Overall | 89.9 | 89.6 | 92.3 | 90.5 |
| Best For | Full-stack coding, Agent workflows, Complex planning | Large codebases, Specs, Careful refactors | Repo edits, Debugging loops, Build/test driven coding | Careful refactors, Long reviews, Spec-heavy coding |
| Weaknesses | Premium pricing, Best results need structured context | Can be cautious, Cost grows on long runs | Needs clean task boundaries, Connector availability matters | Can over-discuss simple patches, Long sessions cost more |