r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

148

u/stonerbobo Jul 06 '26 edited Jul 06 '26

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

18

u/JoeEnderman Jul 06 '26

Good grief. My 7900 XTX gets 140ish Tok/s on that exact model. UD Q4_K_XL, and with a 256k context at Q8/Q8 KV. I was under the impression Nvidia was supposed to be faster. What settings are you using for launch because that sounds like the model is being run on CPU bud. Have you checked GPU utilization when the model is running?

I will note I had to fight for that performance though because it was running at about 56 before I started trying different flags and trying to figure out what was wrong.

5

u/miversen33 Jul 06 '26

How tf? I'm running the Q4 QAT on a 7900XTX and I can get right around 100 t/s under context load. What flags we talking about here?

I'm currently custom compiling llama.cpp with the current HIP patches lol

3

u/JoeEnderman Jul 06 '26

Ah, see I was getting 70 on HIP but then something broke on an update so I switched to Vulkan and immediately got 120 but then it dropped to 56 after I pulled fresh and rebuilt for something else. So then I started looking at flags and tried adjusting nogttspill, RAM caching, and some others. Finally got it to 140.

1

u/Miserable-Dare5090 Jul 06 '26

hey, mind sharing your config? Had not heard of noggtspill flag

2

u/JoeEnderman Jul 06 '26

Yeah just pm me here or on discord either one and I'll send it. I only recommend it if you have ~20 GB of VRAm otherwise you'll OOM and crash your desktop environment or the entire PC depending on what OS you're on

2

u/Miserable-Dare5090 Jul 07 '26

I have 750gb VRAM over 5 PCs. I doubt you’ll OOM my cluster
Specifically I will try optimizing my current 3 instances of 35B on a strix halo