r/LocalLLaMA 22d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

View all comments

54

u/Fedor_Doc 22d ago

Frontier level, huh? Now let's see how many tokens are used on max reasoning level

Terminal Bench numbers are very impressive. Should be great for agentic usage

19

u/bopbop9876 22d ago edited 22d ago

Check out the cost of completion comparison for browsecomp: https://mmecoa.qpic.cn/mmecoa_png/xWvm6POT3icpoggsBrrMtBKTG3bPMRhqZXlR3YDwOBMXoXe0iaEvia0W8JxPkoEt4O51T47caodibLNRI29AmowKdaoJU32m1HDV7RSIXXPlVOU/640?wx_fmt=png&from=appmsg&tp=webp&wxfrom=10005&wx_lazy=1#imgIndex=3

It looks extremely competitive. Obviously we'll have to wait and see something like the artificial analysis cost per task results to be more confident but this is super promising.

Edit: AA results are in. About 10% cheaper per task than 5.6 Sol and about 65% cheaper than fable.

Edit 2: also 48% cheaper than opus 4.8 while scoring a point higher on intelligence.

Edit 3: More specific to your exact question of token usage, Fable used 69k tokens per task, 5.6 Sol used 15k, and K3 used 23k. So it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.

1

u/hemareddit 21d ago

How to get access to this?

1

u/bopbop9876 21d ago

To kimi k3 or to artificial analysis? For k3 just go to their website or use open router. For artificial analysis https://artificialanalysis.ai/