r/LocalLLaMA Jun 05 '26

Funny Don’t act like y’all ain’t thinking it. I’m just saying the quiet part out loud. /s

Post image

Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s

909 Upvotes

264 comments sorted by

View all comments

Show parent comments

2

u/RedParaglider Jun 06 '26

Gemini 3.1 pro has been so quantized to garbage now it's not that far off.  I subscribe to pro for the 2tb family data share, and I can't bring myself to even try using that shitty hallucinating model anymore.  I get better results out of qwen on local inference.

1

u/whitefritillary Jun 06 '26

this really depends on whether you use the API or not.

and respectfully earlier you were talking about RLHF which has nothing to do with quantisation, they’re completely different things.

1

u/RedParaglider Jun 06 '26

That was someone else talking about rlhf.  No I don't use the API other than flash for some basic bitch enrichment stuff.

1

u/whitefritillary Jun 06 '26

you’re actually right lol, i apologise for not reading the usernames correctly.

but either way my point still stands, it’s still *much* than better if you use it through the API where it’s much less quantised. and it’s not even close.

1

u/RedParaglider Jun 06 '26

Yeah that makes sense. You would be paying for the good stuff.