CursorBench has been updated to version 4.0, and Sonnet 5.5, Opus 5.5 and Fable 5.1 are now on the leaderboard. I compared them with the previous generation (Sonnet 5 and Opus 5), looking at cost and tokens per task.
CursorBench 4.0 has a new task set (long-horizon edit, refactor, investigation, intent-understanding, job-management and design-adherence problems), so the same models now score much lower than in my 3.2 post.
The harness matters. I’ve been following my own advice, but the outcome didn’t match the data. But I’m not using Cursor (anymore, sigh).
Using Claude Code, this was my experience, pretty much:

I recently switched to OpenCode, and the difference is noticeable. I’ll tell you about that in another post.
Four cost/score bands
| Band | Target score | Winner | Cost | Why |
|---|---|---|---|---|
| Budget | ≤40% | Sonnet 5.5 (Low–Medium) | $0.50–$0.70 | Cheapest points on the chart; Medium is ~10% better than Low for 1.4× the cost |
| Everyday | 44–48% | Sonnet 5.5 (High) or Opus 5.5 (Low) | $1.67 / $1.17 | Sonnet scores ~9% higher for 1.4× the cost; Opus Low uses 2.4× fewer tokens |
| Mid | 52–56% | Opus 5.5 (Medium–High) | $2.91–$3.97 | High matches Extra High’s score at 57% of the cost |
| Ceiling | 57%+ | Opus 5.5 (Max) | $13.43 | The only tier above 56%, ~3% better than High for 3.4× the cost |
Sonnet 5, Opus 5 and Fable 5.1 win no band.
Cost vs. score graph
Tokens vs. score graph
Takeaways
- The 5.5 generation replaces the 5s. Every Opus 5 and Sonnet 5 tier is beaten on both cost and tokens by some 5.5 tier. Sonnet 5.5 Low outscores Sonnet 5 Max at ~7% of the cost; Opus 5.5 Medium beats Opus 5 Max by ~13% at ~24% of the cost.
- Fable 5.1 doesn’t earn its price here. Its Max tier is the most expensive point on the chart, and Opus 5.5 High scores ~8% higher at ~23% of the cost. Every Fable tier loses on cost; on tokens only Low survives, and Opus 5.5 Medium beats it by ~16% at ~53% of the cost, using ~9% more tokens.
- Opus 5.5 High is the sweet spot. Extra High scores the same for 1.8× the cost and 1.9× the tokens, so it buys nothing.
- Max is poor value on both new models. Opus 5.5 Max is ~3% better than High for 3.4× the cost and 4.1× the tokens. Sonnet 5.5 Max is ~5% better than Extra High for 2.5× the cost and 2.7× the tokens, and it is the most token-hungry point on the chart.
- Sonnet 5.5’s value is between Low and High. High scores ~33% more than Low for 3.3× the cost. Above that it stops making sense: Extra High is only ~1% better than Opus 5.5 Medium, for 1.3× the cost and 2.6× the tokens.
- Opus 5.5’s steep part is Low to Medium. ~20% better for 2.5× the cost, then only ~7% better from Medium to High for 1.4× the cost. Low is still a real option for small tasks.
- Cheaper is not fewer tokens. Sonnet 5.5 High beats Opus 5.5 Low by ~9% for 1.4× the cost, but uses 2.4× the tokens and 1.5× the steps. If those track wall-clock time (I haven’t measured that), Opus 5.5 Low is the faster of the two.
In short:
- Simple, well-scoped edits ⇒ Sonnet 5.5 Low/Medium.
- Everyday work ⇒ Sonnet 5.5 High.
- Feature work, gnarly debugging ⇒ Opus 5.5 High. This is the sweet spot; Medium if you’re counting pennies.
- Need the ceiling? ⇒ Opus 5.5 Max, knowing what it costs.
- Sonnet 5, Opus 5 and Fable 5.1 ⇒ no reason to pick them anymore.
Data is a snapshot of the CursorBench 4.0 leaderboard taken on 2026-09-29; the live page may have since added models or revised numbers (its changelog shows past repricing of existing results). The leaderboard also lists non-Claude models, which I left out. “Tokens” is the leaderboard’s “Tokens / task” column; the page doesn’t say whether that means output tokens or a total. I don’t see error bars on the leaderboard, so treat gaps of a few percent as possibly noise.
Raw data
| Model | Tier | Score % | Cost/task | Tokens/task | Steps/task |
|---|---|---|---|---|---|
| Opus 5.5 | Max | 57.8 | $13.43 | 218,363 | 185 |
| Opus 5.5 | Extra High | 56.0 | $6.98 | 101,083 | 109 |
| Opus 5.5 | High | 56.0 | $3.97 | 53,078 | 68 |
| Opus 5.5 | Medium | 52.5 | $2.91 | 37,954 | 54 |
| Opus 5.5 | Low | 43.7 | $1.17 | 15,811 | 28 |
| Sonnet 5.5 | Max | 55.5 | $9.67 | 271,920 | 170 |
| Sonnet 5.5 | Extra High | 53.1 | $3.88 | 100,158 | 78 |
| Sonnet 5.5 | High | 47.8 | $1.67 | 37,391 | 41 |
| Sonnet 5.5 | Medium | 39.2 | $0.70 | 16,036 | 22 |
| Sonnet 5.5 | Low | 35.8 | $0.50 | 11,668 | 18 |
| Fable 5.1 | Max | 51.8 | $17.28 | 117,236 | 128 |
| Fable 5.1 | Extra High | 51.6 | $13.01 | 87,294 | 101 |
| Fable 5.1 | High | 49.2 | $9.08 | 58,438 | 77 |
| Fable 5.1 | Medium | 46.8 | $7.05 | 45,411 | 63 |
| Fable 5.1 | Low | 45.1 | $5.44 | 34,795 | 51 |
| Opus 5 | Max | 46.6 | $11.95 | 85,384 | 106 |
| Opus 5 | Extra High | 46.1 | $11.43 | 80,094 | 103 |
| Opus 5 | High | 44.7 | $9.00 | 61,405 | 86 |
| Opus 5 | Medium | 43.3 | $6.94 | 45,272 | 72 |
| Opus 5 | Low | 40.7 | $4.87 | 31,995 | 57 |
| Sonnet 5 | Max | 34.1 | $7.17 | 149,257 | 140 |
| Sonnet 5 | Extra High | 32.0 | $4.55 | 83,373 | 102 |
| Sonnet 5 | High | 30.8 | $3.48 | 61,146 | 85 |
| Sonnet 5 | Medium | 28.0 | $2.31 | 39,114 | 65 |
| Sonnet 5 | Low | 24.1 | $1.39 | 23,772 | 46 |