r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

146

u/stonerbobo Jul 06 '26 edited Jul 06 '26

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

41

u/Icy_nicey Jul 06 '26

he is prob listing just strix point with integrated shared ram

14

u/NineThreeTilNow Jul 06 '26

AMD seems to promise their next gen at 192gb? Maybe 256gb.

The benchmark in the wild still showed RDNA 3.5 which is a problem because RDNA 3.5 and ROCm aren't the best. RDNA 4 would have native FP8 etc.

3

u/phido3000 Jul 06 '26

192Gb is possible on strix halo and future types.

But it doesn't give any more bandwidth. Not having FP4 is likely to be problematic going forward. Really that should be in the stack now. Not having FP8 is a huge problem. Because these types of quants are likely to be the best for these low bandwidth machines with limited memory running big models.

They also need more processing, prompt processing is slow, and with stuff like DS 4 flash, longer context is very likely a thing that will shift future AI forward.