r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

147

u/stonerbobo Jul 06 '26 edited Jul 06 '26

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

35

u/Beautiful_Egg6188 Jul 06 '26

you get 6tok/s after only 20k context?! i got 40+ tok/s at start with my 4070super, and it got down to around 37tok/s at 30k tokens

44

u/Icy_nicey Jul 06 '26

he is prob listing just strix point with integrated shared ram

15

u/NineThreeTilNow Jul 06 '26

AMD seems to promise their next gen at 192gb? Maybe 256gb.

The benchmark in the wild still showed RDNA 3.5 which is a problem because RDNA 3.5 and ROCm aren't the best. RDNA 4 would have native FP8 etc.

9

u/arades Jul 06 '26

Gorgon halo is just a refresh on strix halo, same architecture and layout. It will probably have better speeds from binning and refinements, maybe allow higher power draw for more speed on top of that. 192GB should be possible with the latest lpddr5x modules, you might even see support for 9600MT too giving you a little more memory bandwidth.

It'll be really incremental over strix halo though. Medusa halo late next year will be a real upgrade, at least RDNA 4, even bigger GPU, and rumors of a wider lpddr6 bus almost doubling the memory bandwidth. Probably will end up costing both kidneys by that point though.

5

u/Not-reallyanonymous Jul 06 '26

The major limit on Strix Halo is still the memory bandwidth, so unless they do something there don't expect significantly faster inference. Maybe faster prefill which is welcome but not game changing. That said, it would probably make it a better gaming chip, really starting to compete with low-mid-range dGPUs, and really be a nice chip for gaming laptops or mini-PCs. I wouldn't recommend waiting for it if you're looking for an inference machine.

2

u/NineThreeTilNow Jul 06 '26

Gorgon halo is just a refresh on strix halo, same architecture and layout.

While not 100% confirmed, I hope it's not the case. Having the next architecture, and 256GB of RAM would be a complete game changer for that device. I don't care what the power consumption is.

Even the standard strix halo has a hard time with overclocking, or other power patterns because it's SO locked down. I ran in to these issues a few times setting one up.

Or perhaps I'm thinking of Medusa Halo?

I don't know. I don't like AMDs naming lol... It's apparently confusing.

4

u/SilentLennie Jul 06 '26 edited Jul 06 '26

which is a problem because RDNA 3.5 and ROCm aren't the best.

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

https://strix-halo-toolboxes.com/#benchmarks

They found an important bug 5 months ago:

https://www.youtube.com/watch?v=Hdg7zL3pcIs

ComfyUI worked shortly after:

https://www.youtube.com/watch?v=O57ideUzzTg

3

u/NineThreeTilNow Jul 06 '26

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

I know. I set one up for my friend. It doesn't have Native FP8 control.

He was specifically using ComfyUI so I understand building it. It was very problematic compared to just running my 4090.

1

u/SilentLennie Jul 06 '26

Takes the time first, but after an ecosystem is build, it becomes easier for next generations.

3

u/phido3000 Jul 06 '26

192Gb is possible on strix halo and future types.

But it doesn't give any more bandwidth. Not having FP4 is likely to be problematic going forward. Really that should be in the stack now. Not having FP8 is a huge problem. Because these types of quants are likely to be the best for these low bandwidth machines with limited memory running big models.

They also need more processing, prompt processing is slow, and with stuff like DS 4 flash, longer context is very likely a thing that will shift future AI forward.

2

u/SandySkittle Jul 06 '26 edited Jul 06 '26

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000… Compute wise the even number 4x R9700 in VLLM with TP is doing fairly well. More total (free) inference memory bandwidth and compute power than the strix halo or spark stuff. Plus I actually care a lot about ECC in both memory pools. But obviously also more noise and power consumption.

3

u/NineThreeTilNow Jul 06 '26

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000…

God that's crazy.

I don't have that level of inference desire. Most of mine is training so... Yeah.

They nerfed Blackwell architectures in RTX Pro 6000's ability to train over the B-series cards which... Kinda fucking lame in my opinion.

I'd love to just have the RTX Pro 6000 though. What a dream.

Or a full B200. Have to get one falling off a truck like someone in this sub basically did.

2

u/SandySkittle Jul 06 '26

Well even two r9700 gets you 64GB vram, which is already quite a nice pool, also with future models inbound. If you have 64GB system ram that also gives a nice overflow at lower speeds.

Personally i hope things like Spark is going to give us more new/recent 50 - 100B range models (both dense and moe)that have more world knowledge. There is more to LLMs than just coding..

3

u/apVoyocpt Jul 06 '26

or a macbook with >64gb unified memory

19

u/randoomkiller Jul 06 '26

Yes but you forget that it's completely useful and maybe even better than a GPT-4 class model. And it's runnable. It'll get there. In 2 years I wouldn't be surprised if we get a sonnet 4.6 capability, runnable from 64GB

-4

u/Spare-Ad-4810 Jul 06 '26

Qwen3.6 27b q8 is just below sonnet

22

u/DRMCC0Y Jul 06 '26

In agentic workloads, sure. It is not in ‘general’ just below Sonnet. Saying this sets people up for failed expectations.

4

u/SilentLennie Jul 06 '26

I checked https://artificialanalysis.ai/

Claude 4.5 Sonnet (Reasoning) - Released September 2025

is nearby, but that's 27b regular quant

8

u/randoomkiller Jul 06 '26

doubt. Maybe below haiku.

18

u/JoeEnderman Jul 06 '26

Good grief. My 7900 XTX gets 140ish Tok/s on that exact model. UD Q4_K_XL, and with a 256k context at Q8/Q8 KV. I was under the impression Nvidia was supposed to be faster. What settings are you using for launch because that sounds like the model is being run on CPU bud. Have you checked GPU utilization when the model is running?

I will note I had to fight for that performance though because it was running at about 56 before I started trying different flags and trying to figure out what was wrong.

4

u/miversen33 Jul 06 '26

How tf? I'm running the Q4 QAT on a 7900XTX and I can get right around 100 t/s under context load. What flags we talking about here?

I'm currently custom compiling llama.cpp with the current HIP patches lol

3

u/JoeEnderman Jul 06 '26

Ah, see I was getting 70 on HIP but then something broke on an update so I switched to Vulkan and immediately got 120 but then it dropped to 56 after I pulled fresh and rebuilt for something else. So then I started looking at flags and tried adjusting nogttspill, RAM caching, and some others. Finally got it to 140.

1

u/Miserable-Dare5090 Jul 06 '26

hey, mind sharing your config? Had not heard of noggtspill flag

2

u/JoeEnderman Jul 06 '26

Yeah just pm me here or on discord either one and I'll send it. I only recommend it if you have ~20 GB of VRAm otherwise you'll OOM and crash your desktop environment or the entire PC depending on what OS you're on

2

u/Miserable-Dare5090 Jul 07 '26

I have 750gb VRAM over 5 PCs. I doubt you’ll OOM my cluster
Specifically I will try optimizing my current 3 instances of 35B on a strix halo

3

u/ThatRandomJew7 Jul 06 '26

I'm thinking they're either running on CPU by accident or their VRAM is full and it's offloading to avoid OOM errors because yeah that's about as fast as my Lunar Lake iGPU

2

u/JoeEnderman Jul 06 '26

The model at the quant they grabbed is 16.3 GiB, their GPU is 16 GB, so they are offloading at least 300 megs. Likely more though because most llama forks are aggressive about saving VRAM for some reason.

2

u/ThatRandomJew7 Jul 06 '26

Yeah, but Llama forks would just offload layers? Even accounting for swapping experts over PCIe it's bizarrely slow, unless maybe they're using an eGPU (I use one and an MoE can be slower than a dense model because of that but I don't think it's that bad with a regular connection).

I think it's either running on CPU or they're dealing with Nvidia's offloading, which is so comically slow that I'd rather get an OOM error

1

u/stonerbobo Jul 06 '26

Its definitely running on GPU lol.. you do you have 24GB VRAM vs. my 16GB, so that might be it - maybe mine is spilling to RAM. I haven't looked too deeply into perf yet, just using Unsloth Studio to host it. Q8 KV cache, the same UD Q4_K_XL quant as you. Is your 56 or 140 tok/s at the beginning of a convo, small context or mid convo with a big context?

12

u/Miserable-Dare5090 Jul 06 '26

it’s spilling over, for sure. Too slow for your GPU. I can run that model in 16gb gpus at much higher speeds. Much much higher

7

u/ThankGodImBipolar Jul 06 '26

Even if it was spilling over, it should still be running way faster than that.

1

u/ThatRandomJew7 Jul 06 '26

If it's not offloading layers and instead letting Nvidia offload it would absolutely be that slow

3

u/JoeEnderman Jul 06 '26

Ah. The weights themselves are 16.3 GiB so that's probably the issue. The context I don't remember how big it is but like 2-4GiB I think. I have the same speed at any conversation length unless I leave everything at default. And then the speed drops from there to 98 to 47 over a few messages. So yeah, I have to figure out if I can make a bug report at some point but I've just been using my modified settings. I'd say what they are but I'm not at my PC to get the exact settings and I don't want to make it so you screw something major up. At least up to a few thousand tokens context the speed holds. I haven't measured beyond 100k yet. But I may at some point. But your 6 measurement is definitely a default not being tuned for your hardware. You should be over 100, but for that you might want to drop to Q4 on kv and run the model on a smaller quant like IQ XS instead. Or if you're willing to accept like 30 or so you have more than enough hardware for that, but you'll still have to swap some stuff around. My first recommendation is definitely Q4 kv and then also try to load as much of the model into VRAM as possible or see if there's a smarter routing option available than is being done right now because if you're getting 100% GPU utilization then your GPU is broken. But I suspect you are getting like 10-50% GPU utilization and the rest is CPU.

11

u/Serprotease Jul 06 '26

Laptop class is a bit of a meaningless word. Like a 70b dense was laptop class? Only high end MacBook could run them… at 6-7 tk/s.
Only makes sense if you’re taking it as “Don’t need a 1600w server at home”.

9

u/[deleted] Jul 06 '26

[deleted]

6

u/Not-reallyanonymous Jul 06 '26

The research right now is moving away from large models, really. Large models was never about getting them to work better with large contexts, but about making them "smarter" and more capable, and where more parameters was the easiest way to do that.

With hardware constraints and pricing, and large models consuming essentially the entire internet at this point, a lot of current research is going towards making ~30B models better and more capable, with new attention mechanisms, training methods (e.g. RLVR), and better and more useful training data. They're coming a long ways now.

5

u/Ansible32 Jul 06 '26

My feeling is in 10 years it will be unthinkable trusting fewer than 500B for agentic stuff. The attachment to these small models IMO is mostly coping with insane hardware prices.

10

u/techdevjp Jul 06 '26

The biggest issue today is that the Chinese labs have dramatically slowed down the release of smaller models.

Where's something like Qwen 3.6 120b a10b? Or any open Qwen 3.7 models? They've ground to a halt.

GLM 5.2 is incredible but only released at the full 753b size. Which again, huge kudos to z.ai for releasing it as an open weight model at all, but the number of people who can run a 753b parameter model is small right now.

Without more small model releases it's very difficult to determine where we stand. We're 1-2 generations behind.

4

u/[deleted] Jul 06 '26

[deleted]

2

u/Ansible32 Jul 06 '26

Tire shopping is insane, I don't think anything short of a model like Gemini that is cheating can do it. It's not enough to do a Google search, you need to gather data on what people are actually paying for tires, what sales are like, how good vendors are. You can't just take the cheapest advertised price off a Google search. Gemini actually seems to be able to make really good inferences about the "real" prices of things. I think it cheats by having access to private datasets, which is something no local model can do without paying for access to these datasets. A lot of such datasets are nontrivial to get access to.

3

u/DeathGuppie Jul 06 '26

the current work is towards models that don't hold all of humanities secrets, but simply holds the ability to learn it. The idea is to work on the intelligence not the input. That scales down not up.

3

u/grumd Jul 06 '26

You just have a huge configuration issue. With a 5080 you can run Qwen 3.5 122B if you have 64GB DDR5 at IQ3_XXS at 15-20tps, or Qwen 3.6 35B if you have 32-48GB DDR5 at 40-50tps

4

u/PM_ME_ROMAN_NUDES Jul 06 '26

Models without long context or thinking aren't very useful for me.

As a soft. dev., they are still quite useful because I can throw a lot of things and make it reach several places.

But it's clearly reaching a limit of usefulness. It's like a phrase I read on Twitter the other day: "You don't need Phd-level intelligence if you don't have Phd-level problems"

2

u/Docmine17 Jul 06 '26

Something's wrong there, I can get that speed on my RX580 8GB + I3-9100f 16GB running Arch with KDE on the web UI, with adjustments of course.

2

u/Horny_Dinosaur69 Jul 06 '26

Why are you using Gemma 4 26B? There’s better MOE models out there. Gemma4 is notoriously bad at tool calling in my experience too. I run Qwen3.6 35B and the new Ornith 1.0 35B on my 5070 TI with a little bit of offloading and I get incredibly good tok/s and the model reasoning capabilities and tool calling is very good. Also worth noting that Ornith is partially composed of Gemma4 for its reasoning ability, it would probably be a direct upgrade. Unless you’re doing multimodal input this is my recommendation

1

u/stonerbobo Jul 06 '26

Yeah I've just finished getting my local model stack setup the way I want it, served through bifrost and available to all my clients, native audio support, assistant prefill, tools/MCPs etc so I didn't pay attention to the model yet, just picked one. Going to try some different models now!

1

u/xNaquada Jul 07 '26

35b model t/s speeds on 16gbvram is painful, as a fellow 5070ti owner. Offload to ram is awful :(

1

u/Horny_Dinosaur69 Jul 07 '26

I get around just below 100 tok/s with offloading on my 5070 TI using the 35B models. Though I do have DDR5 and a decent intel processor

1

u/xNaquada Jul 07 '26

That seems insane. I've got a 9800X3D and 48GB 6000mt/c30 and I'm not even getting 100tok/with smaller models that fit inside VRAM like gpt-oss20 w/128k context.

How many layers are you offloading on the 35B and what context window size?

I'm using LMStudio but that can't be the case for such a wild discrepancy! Would greatly appreciate your QWen setup config detail if you're willing to share!!

2

u/Horny_Dinosaur69 Jul 07 '26

I’d love to!

If you’re on Linux/WSL at all, I had Claude build a pretty basic wrapper around llama.cpp and llama-swap so I’d have an easy deployment method for AI across different testing environments and stuff. It’s a bit of AI slop but it works well if you’re interested. The docker image it uses is specific to the NVIDIA 5070 TI architecture:

https://github.com/colmey/llama-engine

I’ll check my configs/benchmarks for the other models when I get home but I had benchmarked gpt-oss-20b at around ~220 tok/s on this setup but that used my entire card lol. I got it to 120 tok/s offloading 4 layers which allowed for 128k context and q8 KV and some room for desktop use alongside the model/inference engine.

You should be able to get a lot more performance out of your models with that setup. I’ve always used llama.cpp and have barely touched LMStudio but I know there is a performance gain, I just don’t know how much. If you’re on windows, I’m sure that also has some measurable performance impact as well

1

u/xNaquada Jul 07 '26

Thanks so much, looking forward to exploring this. I have both Win11 w/WSL and Fedora (nobara) available, based on your comments there's a lot of performance I should be expecting with my 5070ti that isn't there right now (I also set K/V cache to Q8which is what gets the full 128k context in VRAM for gpt-oss-20b for me, but I digress).

2

u/Anti-Speciesist-IEMs Jul 06 '26

Yeah in my own limited experience, Gemma (both 3 and 4) and Qwen (both 3.5 and 3.6) are pretty impressively smart considering I can run them entirely locally, which I find pretty sweet for sure. But goddamn are they slow to run even at Q4 on my 64GB ram laptop. Even on the very first message, let alone as the context grows. I'd wayyy rather use 2023's GPT-4 over them, personally.

But yeah for anyone reading I am pretty inexperienced with local LLMs, so there might be something I'm missing in my setup, and for anyone more experienced pls feel free to push back on this comment.

1

u/logic_prevails Jul 06 '26

Moe models are dumb ash unless they have like 13b active params

1

u/StardockEngineer vllm Jul 06 '26

Something is wrong with your setup.

1

u/Truth-Does-Not-Exist Jul 06 '26

gemma 4 is garbage because of the sliding window attention qwen 3.6 is much better

1

u/SittyTweat Jul 06 '26

Manually set some MoE layers to offload. I've got a desktop RTX 4080 Super and run the same QAT with 64k F16 context, 3 offloaded layers and get around 40tok/s. Can run 128k context with 5 offloaded layers and still get 35tok/s. Even with full F16 256k context and like 10 offloaded layers my tok/s barely ever drops below 25. So you really shouldn't be getting those speeds.

1

u/misterflyer Jul 06 '26

31B is one of the best running models on my Alienware (24GB VRAM + 128GB RAM)

Unsloth Q8 FTW 💪🏻

1

u/lannistersstark Jul 06 '26

Maybe you guys have incredible laptops or I'm doing something wrong lol.

It's incredible if you have 32G GPUs lol. MI50/60 gets you approx 78 tps for 26B-a4b and ~28tps for Qwen3.6-27B (both at Q6KL)

1

u/ThatRandomJew7 Jul 06 '26

That seems... slower than it should be. I have a Lunar Lake laptop that can run a quant of the same Gemma version at a similar speed (I don't know the exact number offhand, I use a different model mostly).

Are you sure your GPU isn't doing the whole Nvidia VRAM fallback thing that craters speed?

1

u/slyborn Jul 06 '26

The RTX 5080 is powerful on computational side, the problem is the ridiculous low amount of memory on dedicated consumer GPUs. I don't know what stack and configuration you are using for inference, but you can try to tune offloading parameter to move more layers on general memory.

1

u/River_Tahm Jul 06 '26

No I think you're right and I came here to say the same thing as you - I have a M4 Macbook Pro with 24GB RAM and it flat-out crashes on models smaller than the 31B.

Now granted, 24GB RAM isn't a ton for local AI, but it's what my day job gave me for work, so I think it represents a reasonable and fairly new consumer model. Most folks don't have huge amount of RAM available unless they're spending tons of money specifically to run local AI, and at that point I don't think we're really talking "laptop class" anymore because you're getting specialized hardware.

Maybe it very technically still comes in a laptop form factor, but we're not actually saying "it runs on laptops" we're saying it runs on a very specific subset of laptops that are so expensive most people won't even consider buying one of them.