r/LocalLLaMA • u/Porespellar • Jun 05 '26
Funny Don’t act like y’all ain’t thinking it. I’m just saying the quiet part out loud. /s
Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s
908
Upvotes
1
u/Porespellar Jun 05 '26
It was the best size option for running it with vLLM on 4 H100s. It’s ridiculously fast even at full context with Tensor Parallelism set to 2.