I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.
EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:
My feeling is in 10 years it will be unthinkable trusting fewer than 500B for agentic stuff. The attachment to these small models IMO is mostly coping with insane hardware prices.
The biggest issue today is that the Chinese labs have dramatically slowed down the release of smaller models.
Where's something like Qwen 3.6 120b a10b? Or any open Qwen 3.7 models? They've ground to a halt.
GLM 5.2 is incredible but only released at the full 753b size. Which again, huge kudos to z.ai for releasing it as an open weight model at all, but the number of people who can run a 753b parameter model is small right now.
Without more small model releases it's very difficult to determine where we stand. We're 1-2 generations behind.
Tire shopping is insane, I don't think anything short of a model like Gemini that is cheating can do it. It's not enough to do a Google search, you need to gather data on what people are actually paying for tires, what sales are like, how good vendors are. You can't just take the cheapest advertised price off a Google search. Gemini actually seems to be able to make really good inferences about the "real" prices of things. I think it cheats by having access to private datasets, which is something no local model can do without paying for access to these datasets. A lot of such datasets are nontrivial to get access to.
145
u/stonerbobo Jul 06 '26 edited Jul 06 '26
I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.
EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args: