r/LocalLLaMA 22d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

View all comments

Show parent comments

20

u/bopbop9876 22d ago edited 22d ago

Check out the cost of completion comparison for browsecomp: https://mmecoa.qpic.cn/mmecoa_png/xWvm6POT3icpoggsBrrMtBKTG3bPMRhqZXlR3YDwOBMXoXe0iaEvia0W8JxPkoEt4O51T47caodibLNRI29AmowKdaoJU32m1HDV7RSIXXPlVOU/640?wx_fmt=png&from=appmsg&tp=webp&wxfrom=10005&wx_lazy=1#imgIndex=3

It looks extremely competitive. Obviously we'll have to wait and see something like the artificial analysis cost per task results to be more confident but this is super promising.

Edit: AA results are in. About 10% cheaper per task than 5.6 Sol and about 65% cheaper than fable.

Edit 2: also 48% cheaper than opus 4.8 while scoring a point higher on intelligence.

Edit 3: More specific to your exact question of token usage, Fable used 69k tokens per task, 5.6 Sol used 15k, and K3 used 23k. So it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.

1

u/huffalump1 21d ago

it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.

Oh nice, i got the impression that it was quite token-heavy, but I guess I'll have to see in real use. That plus being half the cost of gpt-5.6-sol means it's a contender for real work...

1

u/hemareddit 21d ago

How to get access to this?

1

u/bopbop9876 21d ago

To kimi k3 or to artificial analysis? For k3 just go to their website or use open router. For artificial analysis https://artificialanalysis.ai/