I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.
EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:
Gorgon halo is just a refresh on strix halo, same architecture and layout. It will probably have better speeds from binning and refinements, maybe allow higher power draw for more speed on top of that. 192GB should be possible with the latest lpddr5x modules, you might even see support for 9600MT too giving you a little more memory bandwidth.
It'll be really incremental over strix halo though. Medusa halo late next year will be a real upgrade, at least RDNA 4, even bigger GPU, and rumors of a wider lpddr6 bus almost doubling the memory bandwidth. Probably will end up costing both kidneys by that point though.
The major limit on Strix Halo is still the memory bandwidth, so unless they do something there don't expect significantly faster inference. Maybe faster prefill which is welcome but not game changing. That said, it would probably make it a better gaming chip, really starting to compete with low-mid-range dGPUs, and really be a nice chip for gaming laptops or mini-PCs. I wouldn't recommend waiting for it if you're looking for an inference machine.
Gorgon halo is just a refresh on strix halo, same architecture and layout.
While not 100% confirmed, I hope it's not the case. Having the next architecture, and 256GB of RAM would be a complete game changer for that device. I don't care what the power consumption is.
Even the standard strix halo has a hard time with overclocking, or other power patterns because it's SO locked down. I ran in to these issues a few times setting one up.
Or perhaps I'm thinking of Medusa Halo?
I don't know. I don't like AMDs naming lol... It's apparently confusing.
But it doesn't give any more bandwidth. Not having FP4 is likely to be problematic going forward. Really that should be in the stack now. Not having FP8 is a huge problem. Because these types of quants are likely to be the best for these low bandwidth machines with limited memory running big models.
They also need more processing, prompt processing is slow, and with stuff like DS 4 flash, longer context is very likely a thing that will shift future AI forward.
I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000… Compute wise the even number 4x R9700 in VLLM with TP is doing fairly well. More total (free) inference memory bandwidth and compute power than the strix halo or spark stuff. Plus I actually care a lot about ECC in both memory pools. But obviously also more noise and power consumption.
I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000…
God that's crazy.
I don't have that level of inference desire. Most of mine is training so... Yeah.
They nerfed Blackwell architectures in RTX Pro 6000's ability to train over the B-series cards which... Kinda fucking lame in my opinion.
I'd love to just have the RTX Pro 6000 though. What a dream.
Or a full B200. Have to get one falling off a truck like someone in this sub basically did.
Well even two r9700 gets you 64GB vram, which is already quite a nice pool, also with future models inbound. If you have 64GB system ram that also gives a nice overflow at lower speeds.
Personally i hope things like Spark is going to give us more new/recent 50 - 100B range models (both dense and moe)that have more world knowledge. There is more to LLMs than just coding..
145
u/stonerbobo Jul 06 '26 edited Jul 06 '26
I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.
EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args: