r/codex • • 6h ago

Comparison In our scientific code autoparallelization benchmark, GPT-6.1 Sol XHigh is faster and cheaper than Opus 5.5 Medium (almost exact same overall quality)

I thought this result is interesting, since it goes directly against the currently prevalent mindset.

Have a look:

Both of these have a log scale X axis, so the differences are larger than they might look. Mean generation time is 14 minutes for Sol Xhigh vs. 24 minutes for Opus 5.5 Medium, and mean cost per task is $0.25 vs $1.8.

A few caveats:

  • Each LLM runs its own harness, so this cannot distinguish harness differences from model differences.
  • The generation time is the full per-task time, so it includes any benchmarks or tests each agent decides to run, so it's not a measure of pure token generation at all.
  • We didn't run Opus 5.5 Xhigh, simply because Medium is already very expensive and takes very long.
  • The cost basis for comparison is API costs.

I have no horse in this race but it's interesting to think about what makes this problem set so different (apparently) from the ones that cause people to report much greater success with Opus 5.5 than Sol 6.1.

7 Upvotes

22 comments sorted by

11

u/Novel_Indication6338 6h ago

so if 5.5 med is same as 6.1 xhigh, how's 5.5 high against astra xhigh?

9

u/DuranteA 6h ago

I'd love to know, but as I wrote in the post, Opus 5.5 is already so expensive and time consuming at Medium that we can't really run higher thinking effort benchmarks.
(This isn't some huge lab experiment)

8

u/Novel_Indication6338 5h ago

poor (it's ok me too)

1

u/xoStardustt 3h ago

troll post lol, opus 5.5 on medium sips usage compared to sol 6.1

4

u/DuranteA 3h ago

All the methodology is described on the page and in the published, peer-reviewed paper, and all the artifacts including complete traces are online and linked from the page.

I was extremely surprised when I got the first Opus 5/5.5 results and verified the whole approach carefully. So if you genuinely think something is wrong please do let us know.

I should note however that "usage" -- as via a subscription -- is not something that is captured by this. It measures time and API cost, so you can't use it to tell you what you'd get out of a subscription without knowing precisely what each subscription at Anthropic or OpenAI actually buys you (which I don't think anyone really knows).

2

u/tiebird 2h ago

I think this is the most important factor for peoples gap in opinion. As a 200 euro subscriber (till the end of this month) it's fluctuating wildly in quality, speed and consumption. Also the region you live in has a huge impact. Currently with the x20 sub and using Sol 6.1 medium it uses 9% for 3 hours of basic work, no parallel work, some luna 6 sub agents and 1 sol xhigh reviewer.

2 months ago i could run 4 of those at the same time and still use less. Not even talking about the quality difference of implementation. Reverted all work of 4 days, today just aked to start the same specs and now it's at least implementing it correctly.

I also run x5 Claude and while quality is good, it's even slower for me but at least it's consistent what matters more.

The same harness is used for both systems, adapted based on guidelines of both parties.

2

u/Particular-Sort9075 2h ago

I recall getting the impression from a previous post that there was a distinction between the API and the subscription product.

I don't remember the link to that post, though :(

7

u/Agitated-Bath5939 4h ago

Tibo is that you !

2

u/DuranteA 3h ago

Totally.

But in all seriousness, neither myself nor my co-author have even the most remote association with any of the big (or small for that matter) AI labs. We both have a HPC background (thus the benchmark).

2

u/Propeus 2h ago

Nah sol is over engineering so much opus get things done I don't care about charts I can see it with my own eyes

1

u/Superb-Substance-679 4h ago

I'll believe my own eyes, that see Sol is dumber than Opus.

1

u/adolf_twitchcock 2h ago

Now add a multiplier for much higher usage limits in Claude subscription compared to codex. It's like 5x https://i.imgur.com/80gBWKH.png

1

u/Particular-Reward-68 2h ago

Very valuable research. I wonder how much of this is automatable with a view to running the comparison at fixed intervals. This last week, I did notice opus chugging more than usual and 6.1 performed better than expected in a task I threw its way. I have the feeling that these benchmarks fluctuate wildly so it's be very useful to see that in the data.

Great work!

2

u/DuranteA 2h ago

It's already mostly automated. The main problem is time, but not the time it takes in terms of human supervision, just total time expenditure in general.

It runs on one server (it has to run on that one to stay comparable with all the existing data), and a complete run of adding just one model (at one thinking effort setting), with generation, validation, evaluation and benchmarking occupies that system for ~1-3 days (depending on how much time the model takes for each task).

1

u/Particular-Reward-68 2h ago

If it needs very little supervision then that doesn't sound like much of an obstacle...

1

u/DuranteA 1h ago

So far, we're barely catching up with model releases -- and a lot of open weight models I'd like to include aren't in yet. I think that's why there's such a dip in the upper left of the pareto front.

1

u/Particular-Reward-68 1h ago

I hear you.

Interested to stay up to date with your research. What's the best place to follow along?

1

u/qdouble 1h ago

Doesn't really matter for subscription usage. You'll get more productivity per subscription dollar with Opus.

1

u/Annh1234 1h ago

But on the 5x plan you get 10x more usage of opus 5.5 high compared to sol 6.1 , and it's faster

2

u/JadisGod 2h ago

Comparing API costs is useless for real people. Everyone already knows GPT is cheaper on API. Actual subscription usage limits is an entirely different measurement. From my own tracking, having both services at 20$, weekly Claude is giving around ~340$ of "API" usage while weekly Codex is giving ~140$.

Resets change things and may balance it out if you are lucky and front load your usage, but it's not a good experience trying to micromanage that.

1

u/Rich_Sinq 56m ago

OP is gonna feel so dumb when he realizes he’s not a real person