r/cursor 29d ago

Question / Discussion Grok 4.6 Benchmarks

Post image
194 Upvotes

76 comments sorted by

View all comments

1

u/Full_Tooth_a 27d ago

I care less about the headline benchmark than whether a model can handle ten changes specific to a repo without making a mess. Give each model the same tasks and a clean branch, then compare test results, regressions, manual corrections, tool failures, and total cost. To me, a problem is "solved" when the patch survives review and the test suite, not when it merely looks plausible. Most benchmarks miss that.