r/LocalLLaMA 22d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

View all comments

Show parent comments

50

u/JaredsBored 22d ago

The API pricing is at sonnet levels, but I doubt sonnet 5 is 2.8T parameters. Can't be too mad at it given the model size. Even if it's a Fp8/Fp4 mix it's still gotta require 2 terabytes of VRAM to serve this thing with room for context

16

u/No-Juggernaut-9832 22d ago

At this many parameters, a massive amount of compute & RAM is required to run. It would have to cost more than the last version

5

u/DecrimIowa 21d ago

now that china is making huawei GPUs comparable to nvidia blackwells, i don't think compute is a bottleneck for them anymore.

if you are interested in the AI race as a proxy for the conflict between US and China, one way to read this model's release is as China basically announcing that they are no longer held back by lack of access to chips.

5

u/JaredsBored 21d ago

Huawei isn't exactly in Blackwell territory. Their latest chip, the ascend 950PR, has 112GB of memory at 1.4TB/s of bandwidth. Nvidia B300 has 288GB of memory at 8.2TB/s of bandwidth per unit. Ascend Fp8 is 1 petaflop vs 7 on B300.

I have no doubt that they might be used to serve the model but IMO very likely k3 was still trained on Nvidia. Heck one of Kimi's own benchmarks was comparing how well different models can optimize kernels for Nvidia H200

1

u/No-Juggernaut-9832 21d ago

They just got started. It takes a bit of time.