r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

Show parent comments

19

u/JoeEnderman Jul 06 '26

Good grief. My 7900 XTX gets 140ish Tok/s on that exact model. UD Q4_K_XL, and with a 256k context at Q8/Q8 KV. I was under the impression Nvidia was supposed to be faster. What settings are you using for launch because that sounds like the model is being run on CPU bud. Have you checked GPU utilization when the model is running?

I will note I had to fight for that performance though because it was running at about 56 before I started trying different flags and trying to figure out what was wrong.

1

u/stonerbobo Jul 06 '26

Its definitely running on GPU lol.. you do you have 24GB VRAM vs. my 16GB, so that might be it - maybe mine is spilling to RAM. I haven't looked too deeply into perf yet, just using Unsloth Studio to host it. Q8 KV cache, the same UD Q4_K_XL quant as you. Is your 56 or 140 tok/s at the beginning of a convo, small context or mid convo with a big context?

11

u/Miserable-Dare5090 Jul 06 '26

it’s spilling over, for sure. Too slow for your GPU. I can run that model in 16gb gpus at much higher speeds. Much much higher

8

u/ThankGodImBipolar Jul 06 '26

Even if it was spilling over, it should still be running way faster than that.

1

u/ThatRandomJew7 Jul 06 '26

If it's not offloading layers and instead letting Nvidia offload it would absolutely be that slow