Agents
Same model, different agent, different result.
Agent wrappers change how well a model reads code, uses tools, recovers from errors and finishes multi-step work.
Agent comparison snapshot
Demo data| Metric | Codex | Claude Code | Cursor Agent | Windsurf |
|---|---|---|---|---|
| Task Success | 94 | 92 | 87 | 84 |
| Codebase Understanding | 93 | 95 | 84 | 82 |
| Tool Use | 96 | 90 | 82 | 80 |
| Autonomy | 91 | 88 | 79 | 78 |
| Recovery after Error | 90 | 91 | 77 | 76 |
| Cost | 78 / Medium | 74 / High | 82 / Medium | 86 / Medium |
| Overall | 92.3 | 90.5 | 82.9 | 81.2 |