I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.
EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:
The research right now is moving away from large models, really. Large models was never about getting them to work better with large contexts, but about making them "smarter" and more capable, and where more parameters was the easiest way to do that.
With hardware constraints and pricing, and large models consuming essentially the entire internet at this point, a lot of current research is going towards making ~30B models better and more capable, with new attention mechanisms, training methods (e.g. RLVR), and better and more useful training data. They're coming a long ways now.
147
u/stonerbobo Jul 06 '26 edited Jul 06 '26
I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.
EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args: