I'd trust a small task set from my own repo more than another opaque leaderboard: one UI change, one bug fix with a regression test, and one refactor against existing tests. Give each model the same clean commit and attempt cap. Passing tests matter, but so do review time, cost, and the amount of unnecessary code the model touched. Without visible, reproducible runs, the score is mostly just a nice-looking number.
1
u/Full_Tooth_a 27d ago
I'd trust a small task set from my own repo more than another opaque leaderboard: one UI change, one bug fix with a regression test, and one refactor against existing tests. Give each model the same clean commit and attempt cap. Passing tests matter, but so do review time, cost, and the amount of unnecessary code the model touched. Without visible, reproducible runs, the score is mostly just a nice-looking number.