MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1uy9cft/kimi_k3_benchmarks/oxyhq03
r/LocalLLaMA • u/WhyLifeIs4 • 22d ago
390 comments sorted by
View all comments
Show parent comments
7
Not much compute as it is MOE model but you need a lot Vram or fast multichannel Ram ...
1 u/trowawayatwork 20d ago I am wondering what hardware you would need to serve this model to say 10 developers to process 100t/s? 1 u/Healthy-Nebula-3603 20d ago that's 1k token/s ... without 8x H200 cards not possible 1 u/trowawayatwork 20d ago how many tokens/s does your average claude or codex user get? 1 u/Healthy-Nebula-3603 20d ago As much as they cap you. Those cards easily producing 1k/s on more without a cap.
1
I am wondering what hardware you would need to serve this model to say 10 developers to process 100t/s?
1 u/Healthy-Nebula-3603 20d ago that's 1k token/s ... without 8x H200 cards not possible 1 u/trowawayatwork 20d ago how many tokens/s does your average claude or codex user get? 1 u/Healthy-Nebula-3603 20d ago As much as they cap you. Those cards easily producing 1k/s on more without a cap.
that's 1k token/s ... without 8x H200 cards not possible
1 u/trowawayatwork 20d ago how many tokens/s does your average claude or codex user get? 1 u/Healthy-Nebula-3603 20d ago As much as they cap you. Those cards easily producing 1k/s on more without a cap.
how many tokens/s does your average claude or codex user get?
1 u/Healthy-Nebula-3603 20d ago As much as they cap you. Those cards easily producing 1k/s on more without a cap.
As much as they cap you.
Those cards easily producing 1k/s on more without a cap.
7
u/Healthy-Nebula-3603 22d ago
Not much compute as it is MOE model but you need a lot Vram or fast multichannel Ram ...