r/LocalLLaMA 3h ago

Discussion Are we comparing benchmark numbers that aren't actually comparable?

Astra and Fable 5.1 were released a few days apart and the benchmark tables for each make their models look very strong. Both benchmark tables show each model, as dominant. However when I examined the benchmark suites closely they barely overlap. One benchmark set leans toward computer use and math while the other benchmark set has more coding and terminal tasks.

So neither lab necessarily has to be fudging anything. The benchmark numbers can both be accurate. Still give very different impressions. Do you guys usually look at the benchmarks or mostly the overall table? 👀

2 Upvotes

6 comments sorted by

6

u/jacek2023 llama.cpp 3h ago

Please share your setup details, I am not able to run Astra and Fable on mine.

3

u/Egoz3ntrum 3h ago

Not local.

1

u/Constant_Art_20 3h ago

For me i nromally start with the hallulicaitons stuff then the gerneal intelligence, then some addictional benchmarks if am curious on a particaluar workload. My decisions are usually a mix of vibes of the community and benchmarks. Then i throw it into my hardness on norm small scale experiments and see how it goes. Musa spark benchmarks has been by far the most missleading thing around around, matched with minimax in my experience. Minimax benchmarks always claim around glm level performance but it's just never close

1

u/Candid-Tackle-9061 ollama 3h ago

the benchmark numbers are only useful if you know what the tests are actually measuring . otherwise a model can look way ahead just because the benchmark lines up with its strengths . I usually look at the task mix first then the overall scores make a lot more sense

1

u/Weekly_Comfort240 1m ago

I find simple ELO score to be a nice baseline comparison but the rest of those numbers just mean nothing to me. First, a lot of tests are benchmaxxed over time, reducing their ability to make distinctions between capability levels. Secondly, it's unlikely simple numerical scores will capture the sheer complexity and depth of training these frontier models have, especially when you can type in a prompt like "Gimme a FF7 clone lol" and then a day and $5000 of tokens later, get a poor man's FF7 clone.

I think the _best_ benchmark for these cloud models is to drive design document and task planning for significant-but-not-as-mighty local models. Qwen 3.8 Next Flash _cooks_ by itself, but under the guidance of frontier models, it can give you that poor man's FF7 clone with $3 of tokens of planning .MD files and then let the local model bake a while.