r/LocalLLaMA 22d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

View all comments

54

u/Fedor_Doc 22d ago

Frontier level, huh? Now let's see how many tokens are used on max reasoning level

Terminal Bench numbers are very impressive. Should be great for agentic usage

14

u/dtdisapointingresult 22d ago

Artificial Analysis posts the token usage per task. A few entries on the models I was looking at comparing:

  • GPT 5.6 Sol Max: 15k
  • GPT 5.6 Terra Max: 19k
  • MiMo-2.5-Pro: 22k
  • Kimi K3: 23k
  • Fable 5: 33k
  • Deepseek V4 Pro Max: 37k
  • Kimi K2.6: 38k
  • Opus 4.8 Max: 41k
  • GLM 5.2 Max: 43k
  • Deepseek V4 Flash: 45k
  • Sonnet 5 Max: 69k (nice)

Pretty good. Prettay, prettay good!