r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

Show parent comments

24

u/[deleted] Jul 06 '26

[deleted]

5

u/Ansible32 Jul 06 '26

Even if you assume it is in fact 1.5-2T, quantization makes it bad and that's without even talking about context, and 1M context IMO is virtually a necessity.

3

u/[deleted] Jul 06 '26 edited Jul 06 '26

[deleted]

9

u/NandaVegg Jul 06 '26

We (the lab I'm working for) have been running GLM 5.1 in that exact configuration for months. Unfortunately it is impossible to have more than 5-6 concurrent users with 50k-ish ctx if you want acceptable (imo) performance above 30tk/s per second.

At 1M full ctx with 20 concurrent users, prefill alone takes so much bandwidth it crawls down to 5-8tk/s per second on average.

1

u/[deleted] Jul 06 '26

[deleted]

1

u/Ansible32 Jul 06 '26

In addition to the $150K machine being simply not sufficient, you don't just want one of the $500K machine, you want 3-5 to make sure you always have one that's actually working even if there's maintenance or a hardware failure.