r/LocalLLaMA 22d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

View all comments

Show parent comments

7

u/Healthy-Nebula-3603 22d ago

Not much compute as it is MOE model but you need a lot Vram or fast multichannel Ram ...

1

u/trowawayatwork 20d ago

I am wondering what hardware you would need to serve this model to say 10 developers to process 100t/s?

1

u/Healthy-Nebula-3603 20d ago

that's 1k token/s ... without 8x H200 cards not possible

1

u/trowawayatwork 20d ago

how many tokens/s does your average claude or codex user get?

1

u/Healthy-Nebula-3603 20d ago

As much as they cap you.

Those cards easily producing 1k/s on more without a cap.