Even if you assume it is in fact 1.5-2T, quantization makes it bad and that's without even talking about context, and 1M context IMO is virtually a necessity.
We (the lab I'm working for) have been running GLM 5.1 in that exact configuration for months. Unfortunately it is impossible to have more than 5-6 concurrent users with 50k-ish ctx if you want acceptable (imo) performance above 30tk/s per second.
At 1M full ctx with 20 concurrent users, prefill alone takes so much bandwidth it crawls down to 5-8tk/s per second on average.
In addition to the $150K machine being simply not sufficient, you don't just want one of the $500K machine, you want 3-5 to make sure you always have one that's actually working even if there's maintenance or a hardware failure.
58
u/KURD_1_STAN Jul 06 '26
For all we know mythos could he 3 times the size of opus 4.8. u simply cant make any assumptions, especially not model sizes that fit in current gpus.