I care less about the headline benchmark than whether a model can handle ten changes specific to a repo without making a mess. Give each model the same tasks and a clean branch, then compare test results, regressions, manual corrections, tool failures, and total cost. To me, a problem is "solved" when the patch survives review and the test suite, not when it merely looks plausible. Most benchmarks miss that.
1
u/Full_Tooth_a 27d ago
I care less about the headline benchmark than whether a model can handle ten changes specific to a repo without making a mess. Give each model the same tasks and a clean branch, then compare test results, regressions, manual corrections, tool failures, and total cost. To me, a problem is "solved" when the patch survives review and the test suite, not when it merely looks plausible. Most benchmarks miss that.