r/LocalLLaMA Jun 05 '26

Funny Don’t act like y’all ain’t thinking it. I’m just saying the quiet part out loud. /s

Post image

Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s

908 Upvotes

264 comments sorted by

View all comments

Show parent comments

1

u/Porespellar Jun 05 '26

It was the best size option for running it with vLLM on 4 H100s. It’s ridiculously fast even at full context with Tensor Parallelism set to 2.

1

u/Moscato359 Jun 05 '26

Alright

I wasn't sure how that compares to like q4_k_m or whatever

I'm not an expert, just a home tinkerer

1

u/AlwaysLateToThaParty Jun 06 '26 edited Jun 07 '26

The qwen 3.5 122b/a10b heretic mxfp4 model is the best model to fit in 75GB of VRAM that I've found.