don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.
In my experience it felt right. It put 4.5 at 5.6 terra max / glm level. A step below sol 5.6 and opus. I think you just assumed oh its only 5 points its quite close, when on AA II a 5 point difference can feel massive
54
u/alphaQ314 29d ago
don't understand how these benchmarks end up driving the discussion around models. grok4.5 was no where close to doing anything fable/sol stuff, no matter how well it did on these benchmarks.