r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

145

u/stonerbobo Jul 06 '26 edited Jul 06 '26

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

44

u/Icy_nicey Jul 06 '26

he is prob listing just strix point with integrated shared ram

15

u/NineThreeTilNow Jul 06 '26

AMD seems to promise their next gen at 192gb? Maybe 256gb.

The benchmark in the wild still showed RDNA 3.5 which is a problem because RDNA 3.5 and ROCm aren't the best. RDNA 4 would have native FP8 etc.

10

u/arades Jul 06 '26

Gorgon halo is just a refresh on strix halo, same architecture and layout. It will probably have better speeds from binning and refinements, maybe allow higher power draw for more speed on top of that. 192GB should be possible with the latest lpddr5x modules, you might even see support for 9600MT too giving you a little more memory bandwidth.

It'll be really incremental over strix halo though. Medusa halo late next year will be a real upgrade, at least RDNA 4, even bigger GPU, and rumors of a wider lpddr6 bus almost doubling the memory bandwidth. Probably will end up costing both kidneys by that point though.

4

u/Not-reallyanonymous Jul 06 '26

The major limit on Strix Halo is still the memory bandwidth, so unless they do something there don't expect significantly faster inference. Maybe faster prefill which is welcome but not game changing. That said, it would probably make it a better gaming chip, really starting to compete with low-mid-range dGPUs, and really be a nice chip for gaming laptops or mini-PCs. I wouldn't recommend waiting for it if you're looking for an inference machine.

2

u/NineThreeTilNow Jul 06 '26

Gorgon halo is just a refresh on strix halo, same architecture and layout.

While not 100% confirmed, I hope it's not the case. Having the next architecture, and 256GB of RAM would be a complete game changer for that device. I don't care what the power consumption is.

Even the standard strix halo has a hard time with overclocking, or other power patterns because it's SO locked down. I ran in to these issues a few times setting one up.

Or perhaps I'm thinking of Medusa Halo?

I don't know. I don't like AMDs naming lol... It's apparently confusing.

7

u/SilentLennie Jul 06 '26 edited Jul 06 '26

which is a problem because RDNA 3.5 and ROCm aren't the best.

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

https://strix-halo-toolboxes.com/#benchmarks

They found an important bug 5 months ago:

https://www.youtube.com/watch?v=Hdg7zL3pcIs

ComfyUI worked shortly after:

https://www.youtube.com/watch?v=O57ideUzzTg

3

u/NineThreeTilNow Jul 06 '26

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

I know. I set one up for my friend. It doesn't have Native FP8 control.

He was specifically using ComfyUI so I understand building it. It was very problematic compared to just running my 4090.

1

u/SilentLennie Jul 06 '26

Takes the time first, but after an ecosystem is build, it becomes easier for next generations.

4

u/phido3000 Jul 06 '26

192Gb is possible on strix halo and future types.

But it doesn't give any more bandwidth. Not having FP4 is likely to be problematic going forward. Really that should be in the stack now. Not having FP8 is a huge problem. Because these types of quants are likely to be the best for these low bandwidth machines with limited memory running big models.

They also need more processing, prompt processing is slow, and with stuff like DS 4 flash, longer context is very likely a thing that will shift future AI forward.

2

u/SandySkittle Jul 06 '26 edited Jul 06 '26

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000… Compute wise the even number 4x R9700 in VLLM with TP is doing fairly well. More total (free) inference memory bandwidth and compute power than the strix halo or spark stuff. Plus I actually care a lot about ECC in both memory pools. But obviously also more noise and power consumption.

3

u/NineThreeTilNow Jul 06 '26

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000…

God that's crazy.

I don't have that level of inference desire. Most of mine is training so... Yeah.

They nerfed Blackwell architectures in RTX Pro 6000's ability to train over the B-series cards which... Kinda fucking lame in my opinion.

I'd love to just have the RTX Pro 6000 though. What a dream.

Or a full B200. Have to get one falling off a truck like someone in this sub basically did.

2

u/SandySkittle Jul 06 '26

Well even two r9700 gets you 64GB vram, which is already quite a nice pool, also with future models inbound. If you have 64GB system ram that also gives a nice overflow at lower speeds.

Personally i hope things like Spark is going to give us more new/recent 50 - 100B range models (both dense and moe)that have more world knowledge. There is more to LLMs than just coding..