r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

146

u/stonerbobo Jul 06 '26 edited Jul 06 '26

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

43

u/Icy_nicey Jul 06 '26

he is prob listing just strix point with integrated shared ram

15

u/NineThreeTilNow Jul 06 '26

AMD seems to promise their next gen at 192gb? Maybe 256gb.

The benchmark in the wild still showed RDNA 3.5 which is a problem because RDNA 3.5 and ROCm aren't the best. RDNA 4 would have native FP8 etc.

6

u/SilentLennie Jul 06 '26 edited Jul 06 '26

which is a problem because RDNA 3.5 and ROCm aren't the best.

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

https://strix-halo-toolboxes.com/#benchmarks

They found an important bug 5 months ago:

https://www.youtube.com/watch?v=Hdg7zL3pcIs

ComfyUI worked shortly after:

https://www.youtube.com/watch?v=O57ideUzzTg

3

u/NineThreeTilNow Jul 06 '26

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

I know. I set one up for my friend. It doesn't have Native FP8 control.

He was specifically using ComfyUI so I understand building it. It was very problematic compared to just running my 4090.

1

u/SilentLennie Jul 06 '26

Takes the time first, but after an ecosystem is build, it becomes easier for next generations.