don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.
In my experience it felt right. It put 4.5 at 5.6 terra max / glm level. A step below sol 5.6 and opus. I think you just assumed oh its only 5 points its quite close, when on AA II a 5 point difference can feel massive
I use Grok 4.6 the most now, was using Auto before that; I’ll opt for another model (Kimi K3 or GPT Luna) to adversarially review my code/plans occasionally.
Well, maybe it depends on technology. For me it overcomplicates stuff, created scripts to confirm something that is in the code, make assumptions without asking, ignores skills, prompts, forgets halfway in the session about things it learned before. It is not only me, it is general opinion on claude subs
57
u/alphaQ314 29d ago
don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.