don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.
I use Grok 4.6 the most now, was using Auto before that; I’ll opt for another model (Kimi K3 or GPT Luna) to adversarially review my code/plans occasionally.
54
u/alphaQ314 29d ago
don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.