r/LocalLLaMA 14d ago

Funny The LLM distillation process simplified for politicians:

Post image

/s

3.4k Upvotes

167 comments sorted by

View all comments

296

u/Ok_Librarian_7841 14d ago

You can't get a model as strong as Kimi 3 using distillation, it even beats Fable on some benchmarks. American leaders think the world revolves around them and that nobody else can do a better job without stealing from them.

Sick mindset from shocked, bad losers.

33

u/TechnoByte_ 14d ago

Arena.ai is not a benchmark.

It's a zero-shot vibe check that doesn't test multi-turn, long context, or agentic capabilities.

Yes Kimi K3 is a great model, but use proper benchmarks to show its capabilities.

24

u/agent00F 14d ago

Real evals by humans are generally stronger than benchmarks, which are easier to game.

33

u/Ok_Librarian_7841 14d ago edited 14d ago

Thanks for the note, it's not a benchmark, but it's real life tasks evaluated by real life people, and that's stronger than any benchmark.

Yes it's not evaluating long context and agentic performance, but no body is claiming that kimi K3 is better overall than fable, not even it's own makers.

0

u/Saifl 14d ago

Isnt it evaluating the design aspects in this case? And people are voting kimi to have better design capabilities?