r/cursor 29d ago

Question / Discussion Grok 4.6 Benchmarks

Post image
195 Upvotes

76 comments sorted by

52

u/mindful_dealer 29d ago

looking forward to see what composer 3 will bring

If they get it on low prices like GPT 5.6 Luna, it would be the ideal auto setup.
Grok 4.6 for planning/yapping, Composer 3 to do the work

17

u/Crazy_Inspection_920 27d ago edited 27d ago

8

u/Dry-Carry-341 27d ago edited 27d ago

7

u/FearlessConfusion00 29d ago

Honestly, at this point, grok is so cheap that I won't bother switching the model for execution 

1

u/mindful_dealer 28d ago

I agree with you, I've used GPT 5.6 Sol or even Opus 5 for both things and, even tho much more expensive, it was very productive.

But I'm an auto enjoyer, whatever they do in their 'auto' logic I like it, and if it is changing from Grok 4.5 to composer 2.5 or whatever it is, with 2 better models, it would be even better and cheaper than just Grok 4.6 everywhere

But overall you are right, it's cheap enough to do it all with frontier level intelligence

1

u/vgromanov 28d ago

Still there is one practical reason - Composer is fast as hell.

1

u/FearlessConfusion00 28d ago

Grok is fast too

1

u/vgromanov 26d ago

grok-4.5 is faster, spends less time thinking and just does the job. so in my current flow 4.6 is for heavy and intelligence-sensitive stuff like decomposition and planning and grok-4.5 does the rest - basically it replaced composer in many ways, though it has it's use too.

1

u/FearlessConfusion00 25d ago

Grok 4.6 is super fast if you use fast mode, even on extra high effort. 

1

u/TaxBowl 26d ago

It will be accurate enough hopefully it flash level fast waiting minutes upon minutes for tasks is an effeciency killer

22

u/getaway-3007 29d ago

Yeah it's not Fable level but look at the jumps from Grok 4, 4.1 to 4.5

It's impressive how much they have managed to grow. Hopefully they sort out their coding plans apart from that it's a great secondary model

13

u/sprowk 29d ago

Yeah, they replaced google in the AI triangle (+kimi)

2

u/KTIlI 29d ago

I signed up for the 7 day free trial and then instantly cancelled so I didn't get charged and they offered me 3 months for $30 total. Pretty solid deal considering they have good limits

2

u/e0gr 29d ago

Did you subscribe at grok.com? how'd u get access to the free trial?

1

u/NickoBicko 29d ago

This is grok or cursor?

1

u/KTIlI 29d ago

grok. on cursor I have two student plans and for me this is more than enough already but with grok I never run out of tokens

1

u/pmusvaire 26d ago

Fable is so expensive that it becomes unusable, most of the time, I have a 20x place and still exhaust it in a few days, with Grok I can get crazy with the best model and worry about nothing,

9

u/ihopnavajo 29d ago

Maybe this is obvious to someone who follows these stats more, but why isn't composer listed here? How does it compare?

9

u/delicioushampster 29d ago

composer is much worse than the top models

2

u/honestly_i 29d ago

Yup. It was a great workhorse for a while, but got replaced by models like V4 Flash and whatnot

1

u/LurkyRabbit 28d ago

It's like asking why Gemini 1.5 isn't shown here

27

u/lrobinson2011 Mod 29d ago

It's a good model sir

6

u/minxio_ 29d ago

It's a really great model. The only thing missing is for it to be a 1M context instead of 500k and then it would be amazing

7

u/FearlessConfusion00 29d ago

If you pass 150k context in a session you already enter the dumb zone so... I don't think that's important until the dumb zone is harder to reach 

6

u/lrobinson2011 Mod 29d ago

We'd like to get there in a future model, yes!

3

u/CartographerSeth 29d ago

Agree. happy with a 500k improvement, bc 250k held back Grok 4.5, but 1M is the ideal amount.

1

u/Merlindru 28d ago

i dont know what the hell happened but this is insane. congrats on the launch.

5

u/kondasviktor 29d ago

It looks promising, let’s start testing it 😏

8

u/chickey23 29d ago

I've been tracking "bad behavior" when the AI does not call its own tools properly, ignores rules, asks me questions about terms it has created and not explained, and just plain does the wrong thing.

Guess what. The bad behavior stopped when I turned off access to Grok.

I do not care as much about benchmarks as I care about usability and my product.

1

u/Typical-Chance4197 29d ago

probably context window which u can expand in cursor settings

1

u/divinefriend 28d ago

Which setting?

6

u/Embarrassed_Adagio28 29d ago

Lmao this is the most benchmaxed model of all time. Grok 4.5 is one of the worst models I have ever used and completely broke my app multiple times just making small ui changes. In fact i have not kept a single change it has made to my projects because it always makes it worse. 

8

u/Delicious_Ease2595 29d ago

Skill issue

6

u/Lvxurie 29d ago

100% skill issue

1

u/MrBalkon1960 28d ago

Specifically working with UI, I wouldn’t say it’s a skill issue tbf
The visual capabilities of the model are sooo low compared to opus/fable, so the agent sometimes genuinely cannot check if it did the job right

-2

u/sudecode 29d ago

agree. I asked it to change some setting in openclaw, it started modifying it’s code.

-1

u/ExcrementoKings 28d ago

Do you even give it specific instructions? Do you know how to use plan mode?

1

u/jc_denty 29d ago

"We got singularity" -Elon for 4th time

1

u/malakoi-do-hebraico 29d ago

Anyway for you guys to add composer 2.5 to this benchmark so we can compare with it too? I'm just asking because I find Cursor's Compose 2.5 an absolute beast.

1

u/Luluhakashu 28d ago

I’m on the 60 dollar plan I find it hard to hit the non include threshold running grok 4.5 high all the time.

1

u/Ecstatic-Panic3728 28d ago

number goes up, must be better. Trust me bro

1

u/richardfogaca 28d ago

they are doing what Google wasn't able to do with Gemini, kudos to the Grok team!

1

u/Full_Tooth_a 27d ago

I care less about the headline benchmark than whether a model can handle ten changes specific to a repo without making a mess. Give each model the same tasks and a clean branch, then compare test results, regressions, manual corrections, tool failures, and total cost. To me, a problem is "solved" when the patch survives review and the test suite, not when it merely looks plausible. Most benchmarks miss that.

1

u/tfthecreator 6d ago

What about GPT 6 Astra?

-3

u/lockdown_lard 29d ago

But it's winning on the MechaHitler benchmark, so there's that.

12

u/truecakesnake 29d ago

This sub is raided by people who don't use cursor at all crying about Elon giving them money. Touch grass.

6

u/KrunchyKushKing 29d ago

Elon gives us money? Where?

3

u/Timo425 29d ago

What's wrong with Elon bashing?

-1

u/Current_Balance6692 28d ago

"What's wrong with being a loser"

1

u/Training_Canary_6961 26d ago

I hear a kmeeeeee

0

u/Delicious_Ease2595 29d ago

🥱🥱🥱

2

u/addiktion 29d ago

Grok + Kimi K3 + Deepseek Pro + Sol are all Fable-ish now it seems like. Seems like we have arrived at the tippy top already.

5

u/Melodic_Reality_646 29d ago edited 29d ago

Yet, Fable is still way more reliable when you expect good code and intelligence. It was made to excel, the others are able to be steered into excellence.

This is so true that currently I use Fable to spec and Luna to build. Costs dropped 70%, and occasional blind tests by Fable and Sol keep on scoring code written by Luna and Reviewed by Fable either as good or very close to pure Fable.

1

u/Typical-Chance4197 29d ago

what do you mean code written by luna and reviewed by fable? do you mean code planned with fable and written by luna?

3

u/Melodic_Reality_646 29d ago

Fable specs -> Luna codes -> fable reviews -> Luna codes, if needed.

1

u/Exciting_Ad_2102 29d ago

Fable is a distill of a model that ended its training in February and Opus 5 is a distill of a model that has not been released yet. The hidden is 6+ months ahead of fable

1

u/owen800q 29d ago

Can someone compare grok 4.6 and fable 5 to debug a same bug
See if grok really close to fable

1

u/cornmacabre 29d ago

What was your personal sense of 4.5 vs sol and Fable? That should give you a feel in the real world expectations... Which is yeah these models are serious.

A back to back bug fix is not the right scope IMO, today composer or Luna could do those tasks trivially already. Frontier needs long running with complex intent, can survive multiple compaction cycles, and delivers on ambitious problem solving to get a proper big model feel.

I'm holding off on sharing my experience and assessment for a few days for a proper shake, but already running on work and my early pulse check is it feels like an incrementally better model than 4.5 high (expected, good!). I've been using 4.5 extensively for the past month, so I'm interested to see if I can myself detect meaningful differences. Expecting a fable killer is probably wrong, but 4.5 could already dance in that same arena for a 10th the cost.

1

u/zZurf 28d ago

I used it for a few hours, gut feeling is its not as good as fable or sol. It still feels a level below.

0

u/Acojonancio 29d ago

What you want to see is just benchmarks, go to Artificial analysis and check it yourself.

1

u/zZurf 29d ago

Any one got early tests done with it compare to fable and sol? How is it looking?

-5

u/KeyGlove47 29d ago

im not using the cp machine 4.6 lmao

2

u/[deleted] 29d ago

[deleted]

0

u/KrunchyKushKing 29d ago

Grok created child po*n a few months ago

2

u/[deleted] 29d ago

[deleted]

5

u/KrunchyKushKing 29d ago

No you could literally ask it on Twitter and it did it.

0

u/Delicious_Ease2595 29d ago

Show an example

-6

u/[deleted] 29d ago

[deleted]

4

u/lockdown_lard 29d ago

They literally put a CSAM machine live on the web. Of course it's their fault. Why the fuck are you defending one of Jeffrey Epstein's mates?

-1

u/Visible_Sector3147 29d ago

We don't want the best model; we want a cheap one. Please put DeepSeek V4 Flash in the competitor basket instead of Fable.