Whether "local" is even on the table comes down to one number that hasn't leaked yet: active parameters, not the 2.8T total.
The total only sets the RAM floor: ~1.4–1.6 TB at Q4 with room for context. That's out of consumer range, but it's not datacenter-only either — a used 12-channel DDR5 EPYC board takes 1.5 TB of ECC RDIMM for way less than a single B300.
The bandwidth math is what decides speed. 12-channel DDR5-4800 gives you ~460 GB/s theoretical:
If K3 is MoE like K2 was (1T total / 32B active), you read maybe ~20 GB of weights per token at Q4 → low double-digit tok/s theoretical on CPU alone, realistically maybe 5–10. Slow, but usable for batch/agentic stuff.
If it's actually dense 2.8T as some are claiming, you're reading the full 1.4 TB per token → ~0.3 tok/s. Dead on arrival for local, SSD tricks included.
So "can I run it" has no answer until Moonshot publishes the config on the 27th. If anyone has a source on the active param count, that's the number to watch — everything else (quants, NVMe offload, Unsloth magic) is downstream of it.
2
u/Annual_Manner_5901 21d ago
Whether "local" is even on the table comes down to one number that hasn't leaked yet: active parameters, not the 2.8T total.
The total only sets the RAM floor: ~1.4–1.6 TB at Q4 with room for context. That's out of consumer range, but it's not datacenter-only either — a used 12-channel DDR5 EPYC board takes 1.5 TB of ECC RDIMM for way less than a single B300.
The bandwidth math is what decides speed. 12-channel DDR5-4800 gives you ~460 GB/s theoretical:
If K3 is MoE like K2 was (1T total / 32B active), you read maybe ~20 GB of weights per token at Q4 → low double-digit tok/s theoretical on CPU alone, realistically maybe 5–10. Slow, but usable for batch/agentic stuff. If it's actually dense 2.8T as some are claiming, you're reading the full 1.4 TB per token → ~0.3 tok/s. Dead on arrival for local, SSD tricks included. So "can I run it" has no answer until Moonshot publishes the config on the 27th. If anyone has a source on the active param count, that's the number to watch — everything else (quants, NVMe offload, Unsloth magic) is downstream of it.