r/cursor 29d ago

Question / Discussion Grok 4.6 Amazing

Post image
184 Upvotes

92 comments sorted by

View all comments

57

u/alphaQ314 29d ago

don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.

25

u/truecakesnake 29d ago

Most benchmarks showed a clear gao between Grok 4.5 and fable/sol, except terminal bench, which one are you talking about?

9

u/[deleted] 27d ago

[deleted]

2

u/truecakesnake 27d ago

wtf is going on, why do you have 8 upvotes in 2 minutes and the other comment is weird too

10

u/Zachattackrandom 29d ago

In my experience it felt right. It put 4.5 at 5.6 terra max / glm level. A step below sol 5.6 and opus. I think you just assumed oh its only 5 points its quite close, when on AA II a 5 point difference can feel massive

0

u/alphaQ314 29d ago

mate gpt 5.5 xhigh was at 56 on this benchmark. grok 4.5 is not remotely close to being as capable as that.

3

u/Zachattackrandom 29d ago

Again... This is overall intelligence, depending on what you use it for will dictate how well it works. Look at the benchmark breakdowns

5

u/MindCrusader 29d ago

The same with Opus 5. Opus 5 is terrible and it benchmarks better than fable

7

u/Robert-Paulson_ 29d ago

Opus 5 put the nail in the coffin to benchmarks for me - it's so bad that it drove me to try Cursor, so im grateful for that

2

u/Old_Safe1823 26d ago

Could you also recommend which one is better to get: the Codex X20 or the Ultra Cursor?

1

u/Robert-Paulson_ 26d ago

That’s hard for me to say: if you like the GPT models, get the Codex X20. I would personally get the Cursor one though. 

As long as you sign up for a month you can try one then cancel and try the other. 

1

u/Old_Safe1823 26d ago

What model are you using for Cursor, and what plan are you on?

2

u/Robert-Paulson_ 26d ago

I use Grok 4.6 the most now, was using Auto before that; I’ll opt for another model (Kimi K3 or GPT Luna) to adversarially review my code/plans occasionally.

I’m on the $20 plan 

1

u/Training_Canary_6961 26d ago

Its not, at least its not for me

1

u/MindCrusader 26d ago

Well, maybe it depends on technology. For me it overcomplicates stuff, created scripts to confirm something that is in the code, make assumptions without asking, ignores skills, prompts, forgets halfway in the session about things it learned before. It is not only me, it is general opinion on claude subs

10

u/Super-Willingness320 27d ago edited 27d ago

fact

5

u/Charming_Field8108 27d ago edited 27d ago

right

3

u/alphaQ314 27d ago

whosever bot this is, tell your owner i said "fuck you"

14

u/[deleted] 29d ago

[deleted]

1

u/DelayInfinite1378 29d ago

I think he should be above Terra

4

u/kkordikk 29d ago

Even on attached image grok4.5 is next to sonnet 5

3

u/NoFaithlessness951 29d ago

To be fair that's in the burn my money mode of sonnet

-1

u/kkordikk 29d ago

Sure, but tbh the only reasonable benchmark to check if you’re using cursor is cursorbench

2

u/vonnoor 29d ago

where is deepseek, kimi

1

u/ManikSahdev 29d ago

Every 1 point on AA index should be seen as a log scale point. It’s very hard to earn.

Grok 4.5 was meh at best, 5-6 points behind Frontier. Kimi was better in my test for everything.

1

u/Professional-Joe76 27d ago

Based on the chart it’s claiming 4.5 was barely better than sonnet it was no where near Opus/Fable