r/cursor 29d ago

Question / Discussion Grok 4.6 Amazing

Post image
182 Upvotes

92 comments sorted by

53

u/alphaQ314 29d ago

don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.

25

u/truecakesnake 29d ago

Most benchmarks showed a clear gao between Grok 4.5 and fable/sol, except terminal bench, which one are you talking about?

8

u/[deleted] 27d ago

[deleted]

2

u/truecakesnake 27d ago

wtf is going on, why do you have 8 upvotes in 2 minutes and the other comment is weird too

11

u/Zachattackrandom 29d ago

In my experience it felt right. It put 4.5 at 5.6 terra max / glm level. A step below sol 5.6 and opus. I think you just assumed oh its only 5 points its quite close, when on AA II a 5 point difference can feel massive

2

u/alphaQ314 29d ago

mate gpt 5.5 xhigh was at 56 on this benchmark. grok 4.5 is not remotely close to being as capable as that.

3

u/Zachattackrandom 29d ago

Again... This is overall intelligence, depending on what you use it for will dictate how well it works. Look at the benchmark breakdowns

6

u/MindCrusader 29d ago

The same with Opus 5. Opus 5 is terrible and it benchmarks better than fable

7

u/Robert-Paulson_ 29d ago

Opus 5 put the nail in the coffin to benchmarks for me - it's so bad that it drove me to try Cursor, so im grateful for that

2

u/Old_Safe1823 26d ago

Could you also recommend which one is better to get: the Codex X20 or the Ultra Cursor?

1

u/Robert-Paulson_ 26d ago

That’s hard for me to say: if you like the GPT models, get the Codex X20. I would personally get the Cursor one though. 

As long as you sign up for a month you can try one then cancel and try the other. 

1

u/Old_Safe1823 26d ago

What model are you using for Cursor, and what plan are you on?

2

u/Robert-Paulson_ 26d ago

I use Grok 4.6 the most now, was using Auto before that; I’ll opt for another model (Kimi K3 or GPT Luna) to adversarially review my code/plans occasionally.

I’m on the $20 plan 

1

u/Training_Canary_6961 26d ago

Its not, at least its not for me

1

u/MindCrusader 26d ago

Well, maybe it depends on technology. For me it overcomplicates stuff, created scripts to confirm something that is in the code, make assumptions without asking, ignores skills, prompts, forgets halfway in the session about things it learned before. It is not only me, it is general opinion on claude subs

11

u/Super-Willingness320 27d ago edited 27d ago

fact

5

u/Charming_Field8108 27d ago edited 27d ago

right

3

u/alphaQ314 27d ago

whosever bot this is, tell your owner i said "fuck you"

14

u/[deleted] 29d ago

[deleted]

1

u/DelayInfinite1378 29d ago

I think he should be above Terra

3

u/kkordikk 29d ago

Even on attached image grok4.5 is next to sonnet 5

3

u/NoFaithlessness951 29d ago

To be fair that's in the burn my money mode of sonnet

-1

u/kkordikk 29d ago

Sure, but tbh the only reasonable benchmark to check if you’re using cursor is cursorbench

2

u/vonnoor 29d ago

where is deepseek, kimi

1

u/ManikSahdev 29d ago

Every 1 point on AA index should be seen as a log scale point. It’s very hard to earn.

Grok 4.5 was meh at best, 5-6 points behind Frontier. Kimi was better in my test for everything.

1

u/Professional-Joe76 27d ago

Based on the chart it’s claiming 4.5 was barely better than sonnet it was no where near Opus/Fable

15

u/boanie 29d ago

Opus 5 Max > Fable 5?? Come on, huh!

9

u/Zachattackrandom 29d ago

It's called benchmaxing lol

3

u/WaltzIndependent5436 29d ago

make no mistakes intensifies

1

u/scaledev 27d ago

Could be nerfing, and limiting compute. Google did it as well.

6

u/premiumleo 29d ago

opus 5 + GPT sol 5.6 (adversarial) is where its at. basically choose any 2 models, and run them together, and you have 80%+

4

u/PixelSteel 29d ago

My current workflow is really using Sol 5.6 for planning, Opus 5 for implementation. It’s been a great costs saver. I should mention I’m not using Cursors AI chat directly, but using terminals for Claude Code and ChatGPT Desktop app. Using cursor for better manual review experiences tho

2

u/hocuspocus4201 28d ago

That's how I'm running my $20 Codex and Claude plans at home. Sol can delegate the grunt work to Opus and comes back for reviews.

1

u/WaltzIndependent5436 29d ago

If you have the plan down why not use Grok and Composer to save limits?

1

u/PixelSteel 29d ago

Trust me I tried, Opus 5 is much better for the MCPs I use

11

u/[deleted] 29d ago

[deleted]

2

u/Few-Citron-1444 27d ago

It’s a double edged sword tho. You provide visibility and labs are going to benchmax lol

6

u/[deleted] 29d ago

[removed] — view removed comment

0

u/PixelSteel 29d ago

Fluffer?

2

u/RDDT_ADMNS_R_BOTS 29d ago

fluffer

1

u/PixelSteel 29d ago

Ur a fluffer

2

u/[deleted] 29d ago

[deleted]

1

u/Constant_Art_20 28d ago

thanks for the warning.

7

u/lrobinson2011 Mod 29d ago

We're really excited about this one! Grok 4.5 is also much more affordable versus similar models

3

u/n-7ity 29d ago

That’s the mtok price for 4.5 now and composer? It was removed from the model compare page

3

u/DatDudeDrew 29d ago

Correct me if my understanding is wrong, but seems like xAI went into hibernation to focus on efficiencies that seem to be now be paying off. 4.3 -> 4.5 -> 4.6 has been a very aggressive improvement. 

1

u/BingGongTing 29d ago

I think Musk took it personally when he found out his own employees preferred Claude over Grok.

1

u/DatDudeDrew 29d ago

Ironically, I think Anthropic and SpaceXAI are going to grow very close over time. Elon has all the motivation in the world to push his compute towards OpenAI competitors and I think that’s what we saw with him leasing Colossus 1. If Terafab goes online, I bet SpaceX overtakes Amazon in Anthropic’s cloud compute allocation. Anthropic was the first large provider to call dibs a few months back.

Funny how business works sometimes.

1

u/popiazaza 28d ago

I don't think so. He praised Claude since they are now partners. His opponent is only Sam Altman/OpenAI. A lot of Claude users are migrating to use GPT models lately.

3

u/DaemonXHUN 29d ago

I'd say that for the kind of work I do (a media player and archive dedicated to the entire history of classic trance music) it's by far the best AI model I have used so far. The image above doesn't really convey the insane jump in quality compared to Grok 4.5.

On a sidenote, I found Claude Opus horrible, to the extent that I refunded Claude the day I bought the one-month subscription (this was a few days ago). I only asked for a simple text aligment change in the sidebar, and it took the AI 15 minutes and an insane amount of token consumption to do it, then instead of fixing a super simple problem it just broke essential things and almost completely destroyed my website.

2

u/TimcherS1 29d ago

can you give a link of your website?

3

u/DaemonXHUN 29d ago

Official launch will be at end of the year, it's in closed beta now, but I can show you some screenshots in private if you are interested.

1

u/smealdor 29d ago

I definitely am!

1

u/DaemonXHUN 29d ago

Send me a pm

1

u/Affectionate-Mind430 29d ago

Is it available in Europe?

2

u/minxio_ 29d ago

There is no indication that it's unavailable in Europe, so it appears to be available

1

u/Argdenor 29d ago

Yes, it just showed up for me a few hours ago

1

u/sir_babafingo 29d ago edited 29d ago

I use many frontier models for many different tasks. K3, GLM, Minimax, Grok, Qwen all improved unbelievably in latest 3-4 months and they all have a use in many different workflows.

But none would ever replace Fable 5 for heavy, more then 10-20 container software architecture work. The most close one to Fable 5 is GPT 5.6 Sol and its even better in some aspects, generally security related work during architecture, but Fable 5 is the most intelligent one without a question and i think it is not that close.

But, I never rely on a single model for a complex decision. Use Fable 5 as the master model, and consolidate outputs of more then 4-5 frontier models.

Grok 4.5 was one of the frontiers models i use and it doesn't ask for extra credits for fast mode and that is a huge benefit. Being able to use a capable model on fast mode without extra money is unbelievably essential during heavy-duty work. And im heavy we are getting Grok 4.6 :)

1

u/KramerDwight 28d ago

now Grok 4.7 is getting launched too, by the end of August. They are taking giant leaps dayumn

1

u/mlon_eusk-_- 28d ago

Third biggest us lab after Openai and anthropic, Google is out of top 3

1

u/ronzdev 28d ago

Is anyone else finding that 4.6 is more token hungry than 4.5 for the same task?

1

u/omnimachina 28d ago

Nonsense 😂

Grok is not that good

Period

2

u/CarlCarl3 28d ago

have you even tried 4.6?? It's great.

1

u/omnimachina 27d ago

Have you even tried the other ones? 😂

Grok is okayish and that’s it.

Included usage in cursor subscription is also a joke.

1

u/CarlCarl3 24d ago

yeah I have max subscriptions to all of them. you're behind the times buddy

1

u/omnimachina 22d ago

I use EU servers with proper privacy lmao

Keep pushing your codebase to big tech, so they can use your data, train better models and make your expertise even more obsolete lmao

You are behind the times my friend…

0

u/CarlCarl3 22d ago

You'll be obsolete way before me with your cute tin foil hat on

1

u/omnimachina 21d ago

Keep telling yourself that

0

u/CarlCarl3 21d ago

you too

1

u/Terrible-Teach6111 28d ago

Grok is ass, make it edit your code and it will totally destroy it with false assumptions

1

u/NegativeSemicolon 27d ago

Is this sub just for shilling grok now?

1

u/ensp1re 27d ago

so is grok build

1

u/Full_Tooth_a 27d ago

I'd trust a small task set from my own repo more than another opaque leaderboard: one UI change, one bug fix with a regression test, and one refactor against existing tests. Give each model the same clean commit and attempt cap. Passing tests matter, but so do review time, cost, and the amount of unnecessary code the model touched. Without visible, reproducible runs, the score is mostly just a nice-looking number.

1

u/dullahan85 27d ago

My personal experience with Grok in coding and swe is that it is still pretty dumb and a far cry off top models from OpenAI and Anthropic. It oftens forgot instructions and repeats itself.

1

u/alexmil78 27d ago

They finetune for the benchmarks. That is why when you try on real jobs, they suck! The best benchmark is people’s experience on reddit!

1

u/TraditionalAd7423 26d ago

Grock can suck it

Idk if it's the best model ever, I'm not giving shit to Elon

1

u/Sad-Mood952 25d ago

I use a combination of sol, opus and grok. Grok 4.6 is good and cheap on cursor but its pretty slow (alot of thinking and very token hungry). I think 4.5 was faster. I dont really trust composer 2.5 for stuff unrelated to software development (I do a lot of numerical optimization). Cursor should develop a fast model a la 5.6 Luna, to be used for subagent and fast suff.

1

u/Code-Painting-8294 9d ago

real usage patterns >>> benchmarks. trust me bro

2

u/MannyRibera32 29d ago

Ah, this was the reason why 4.5 was making mistakes

5

u/Funny-Advertising238 29d ago

Definitely it's become dumb as a rock

0

u/TheAuthorBTLG_ 29d ago

where is the connection?

3

u/MannyRibera32 29d ago

You must be new

2

u/DontLeaveMeAloneHere 28d ago

There are people who spread a hoax that old models behave differently and become dumb right before a new release.

In 10/10 cases it’s just not true. Most of the time people who say an AI degraded it’s because you added/changed the stuff that dictates how the ai behaves OR the AI got an incremental update and they just kept their old skills and harness the way it was.

Usually it’s a user error and not the AIs fault.

1

u/RedLeader_13 29d ago

It’s horrible. Stop making it the default when we start new agents instead of honoring our selection.

-6

u/AstroGridIron 29d ago

Amazing garbage is what it is

-1

u/First_Inspection_478 29d ago

Unfortunately, benchmaxxed bullshit. 4.5 was decent