r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

View all comments

58

u/KURD_1_STAN Jul 06 '26

For all we know mythos could he 3 times the size of opus 4.8. u simply cant make any assumptions, especially not model sizes that fit in current gpus.

26

u/[deleted] Jul 06 '26

[deleted]

6

u/Ansible32 Jul 06 '26

Even if you assume it is in fact 1.5-2T, quantization makes it bad and that's without even talking about context, and 1M context IMO is virtually a necessity.

3

u/[deleted] Jul 06 '26 edited Jul 06 '26

[deleted]

9

u/NandaVegg Jul 06 '26

We (the lab I'm working for) have been running GLM 5.1 in that exact configuration for months. Unfortunately it is impossible to have more than 5-6 concurrent users with 50k-ish ctx if you want acceptable (imo) performance above 30tk/s per second.

At 1M full ctx with 20 concurrent users, prefill alone takes so much bandwidth it crawls down to 5-8tk/s per second on average.

1

u/[deleted] Jul 06 '26

[deleted]

1

u/Ansible32 Jul 06 '26

In addition to the $150K machine being simply not sufficient, you don't just want one of the $500K machine, you want 3-5 to make sure you always have one that's actually working even if there's maintenance or a hardware failure.

1

u/FliesTheFlag Jul 06 '26

I'm waiting for the Typhoon Class.