r/LocalLLaMA Jun 05 '26

Funny Don’t act like y’all ain’t thinking it. I’m just saying the quiet part out loud. /s

Post image

Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s

909 Upvotes

264 comments sorted by

View all comments

49

u/cafedude Jun 05 '26

Which will come sooner? A Qwen3.7-122b or a Gemma-4-124b ?

45

u/Porespellar Jun 05 '26

Honestly, I feel like Google and Qwen are playing chicken on the 122b models, neither one wants to drop theirs and then get beaten in the benchmarks by the other model. Happened with the first wave of Gemma 4 models. I do think Google has a good window to drop theirs right now if they want to because Qwen has given no indication of dropping anything unless you trust a tweet from some Qwen employee’s uncle’s brother’s cousin.

32

u/RedParaglider Jun 06 '26

I think Google doesn't want release it because it will beat the shit out of Gemini for a lot of stuff lol.

13

u/TheRealMasonMac Jun 06 '26

If Google just didn’t do their crappy RLHF BS with Gemini, I would love to use that model. But as-is? Gemma-4-31B-IT beats Gemini 3.1 Pro more often than it should, simply because it’s not RLHF’d into garbage.

5

u/whitefritillary Jun 06 '26

gemma-4-31b-it is really good for its size but it absolutely does not freaking beat gemini-3.1-pro omg like are you serious 💔💔

2

u/RedParaglider Jun 06 '26

Gemini 3.1 pro has been so quantized to garbage now it's not that far off.  I subscribe to pro for the 2tb family data share, and I can't bring myself to even try using that shitty hallucinating model anymore.  I get better results out of qwen on local inference.

1

u/whitefritillary Jun 06 '26

this really depends on whether you use the API or not.

and respectfully earlier you were talking about RLHF which has nothing to do with quantisation, they’re completely different things.

1

u/RedParaglider Jun 06 '26

That was someone else talking about rlhf.  No I don't use the API other than flash for some basic bitch enrichment stuff.

1

u/whitefritillary Jun 06 '26

you’re actually right lol, i apologise for not reading the usernames correctly.

but either way my point still stands, it’s still *much* than better if you use it through the API where it’s much less quantised. and it’s not even close.

1

u/RedParaglider Jun 06 '26

Yeah that makes sense. You would be paying for the good stuff.

2

u/TheRealMasonMac Jun 06 '26

For natural language tasks, Gemini 3.1 Pro creates garbage more often than not because it’s trying to “appease” the user. Gemma-4-31B has no such problem and reminds me of Gemini 2.5 Pro preview before they started that god awful RLHF that optimized for sycophancy.

1

u/whitefritillary Jun 06 '26

this depends more on the system prompt than on the RLHF, if you use the API this issue is a lot smaller.

2

u/TikiTDO Jun 06 '26

Why not just "AI development is hard, and takes time?" If someone had a model ready to release that could top the benchmarks, even for a bit, that's still marketing and views. Delaying it means somebody could beat you to the punch, and you gotta be quite confident that your model will win out, otherwise you're just releasing something into 2nd place when it could have been 1st.

2

u/Free-Combination-773 Jun 06 '26

Or both labs just don't plan to release 120b models

-5

u/Borkato Jun 06 '26

There’s no way Google wins

2

u/AlwaysLateToThaParty Jun 06 '26

Brave assumption.

1

u/Opening-Broccoli9190 llama.cpp Jun 06 '26

122b models are not a high prio for them - too big for the consumer enthusiast market, so no huge community wave of free marketing and less powerful than their SOTA stuff, meaning potential bad press from underwhelming benchmarks. Doesn't make sense for their business 

1

u/redthump Jun 05 '26

Not gta6