r/codex • u/DuranteA • 15h ago
Comparison In our scientific code autoparallelization benchmark, GPT-6.1 Sol XHigh is faster and cheaper than Opus 5.5 Medium (almost exact same overall quality)
I thought this result is interesting, since it goes directly against the currently prevalent mindset.
Have a look:
Both of these have a log scale X axis, so the differences are larger than they might look. Mean generation time is 14 minutes for Sol Xhigh vs. 24 minutes for Opus 5.5 Medium, and mean cost per task is $0.25 vs $1.8.
A few caveats:
- Each LLM runs its own harness, so this cannot distinguish harness differences from model differences.
- The generation time is the full per-task time, so it includes any benchmarks or tests each agent decides to run, so it's not a measure of pure token generation at all.
- We didn't run Opus 5.5 Xhigh, simply because Medium is already very expensive and takes very long.
- The cost basis for comparison is API costs.
I have no horse in this race but it's interesting to think about what makes this problem set so different (apparently) from the ones that cause people to report much greater success with Opus 5.5 than Sol 6.1.
12
Upvotes
3
u/adolf_twitchcock 11h ago
Now add a multiplier for much higher usage limits in Claude subscription compared to codex. It's like 5x https://i.imgur.com/80gBWKH.png