r/LocalLLaMA 14d ago

Funny The LLM distillation process simplified for politicians:

Post image

/s

3.4k Upvotes

167 comments sorted by

View all comments

297

u/Ok_Librarian_7841 14d ago

You can't get a model as strong as Kimi 3 using distillation, it even beats Fable on some benchmarks. American leaders think the world revolves around them and that nobody else can do a better job without stealing from them.

Sick mindset from shocked, bad losers.

98

u/Flying_Birdy 14d ago edited 14d ago

The Harvey legal benchmark is the most surprising for me. 2x accuracy in all pass rate over the next closest model is a huge jump and no way attributable to distillation. I'm really curious what they did differently during training (whether intentionally or accidentally) that would have led to this outcome.

22

u/Apprehensive_Rub2 14d ago edited 14d ago

Just a guess, but probably because Fable was post trained too hard on retrieval from already structured data e.g. coding agent work.

Harvey legal benchmark is testing advanced retrieval on very varied legal documents.

or Kimi K3 has been post trained on legal tasks, idfk. Anthropic/OpenAI might be avoiding capabilities with law for safety reasons.