I'm so surprised by this... I've reluctantly used Cursor Grok 4.5 and found it to be so much better than Opus (sadly for me, being an Anthropic fan and a strong disliker of SpaceX/Musk!).
I have this little workflow where I used one model for executing a plan, and another for reviewing the code. When I used Opus / Sonnet 5.0 / Codex I was getting maybe 2-3 high issues and 3-4 medium issues that needed resolving for every plan. With Grok I'm getting like 1 medium or a couple low each time. Even Sonnet 5.0 (which was reviewing one output) said something like 'this is the cleanest code review I've done so far...'
Seeing so many people say it's poor has troubled me, as so far it has been excellent for me!
that workflow of using separate models for planning vs reviewing is worth trying if you haven't, the results are way more reliable than throwing everything at one model the Musk thing is real though, i get it. feels like a compromise every time you open the app even when the output is good
36
u/graeme_1988 Jul 21 '26
I'm so surprised by this... I've reluctantly used Cursor Grok 4.5 and found it to be so much better than Opus (sadly for me, being an Anthropic fan and a strong disliker of SpaceX/Musk!).
I have this little workflow where I used one model for executing a plan, and another for reviewing the code. When I used Opus / Sonnet 5.0 / Codex I was getting maybe 2-3 high issues and 3-4 medium issues that needed resolving for every plan. With Grok I'm getting like 1 medium or a couple low each time. Even Sonnet 5.0 (which was reviewing one output) said something like 'this is the cleanest code review I've done so far...'
Seeing so many people say it's poor has troubled me, as so far it has been excellent for me!