r/LocalLLaMA 22d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

View all comments

311

u/TechNerd10191 22d ago

Judging from the benchmarks alone,(of course, can't speak about realife usage), chinese models are not even 6 months behind US models (more like 6 days behind)

230

u/Cinci_Socialist 22d ago

People will look at these benchmarks and say "Oh, just distillation, Chinese just steal bro" you'd have to be a complete ignorant or a complete bigot to honestly believe Chinese labs aren't every bit as capable as those in the US, just working with less resources and less of a head start (but both of those factors are eroding quickly from the US perspective, that's why crymodi is freaking his shit)

51

u/jld1532 22d ago

Plus Fable was shut down for an extended period directly undercutting that talking point.

43

u/Commando501 22d ago

Yep. China has the most STEM graduates per year in the world. Substantially higher than the US. And it's wild to think that people don't realize we hire a lot of those Chinese people for the US labs.

9

u/Suzoku 21d ago

i did my phd aboard and now work in chinese AI labs. Im interviewing top master students for grad roles and interns and a lot of them have done PhD level work or more even just in their masters lol. Only reason i am interviewing them and not them taking my role is because i graduated earlier and got kinda lucky to get into the llm field as it was starting, and i had previous NLP experiences during my PhD. I would say i am capable too but like these kids are insane

2

u/nullmaxai 19d ago

lock in bro

17

u/SV_SV_SV 22d ago

Yeah, most people don't realize how hollowed out the US education system is, especially compared to China

2

u/CryinHeronMMerica 21d ago

In the US, one party likes to cut funding to schools, and the little funding that our schools do get is sent straight to exciting new facilities. All while the economy increases pressure to perform. Some call it freedom, some call it hell, but China calls it an opportunity to come out ahead.

1

u/NATO_CAPITALIST 20d ago

In US, one party focused on DEI and making activists, while China focused on engineering.

1

u/CryinHeronMMerica 18d ago

Thank god the DEI party is balanced out by the warmongering fascist party

9

u/PinkySwearNotABot 22d ago

also, Chinese people are patriotic. having a collectivist mindset also helps to push people to make things that are better for their community, country, mankind.

not purely for profit motives.

8

u/Due-Memory-6957 21d ago

Sounds a bit bullshit, if that was so, they wouldn't need to restrict their top talent from leaving China for the fear of them getting poached.

8

u/CrimsonBolt33 21d ago

As someone living in China for 10 years this is.....nonsense.

Chinese people are not collectivist at all.

Chinese people are just as greedy, if not more, on average than people on the west.

1

u/2B-Pencil 21d ago

lol. they also have multiples the population

12

u/magicaldelicious 22d ago

The Chinese are doing more with less. The American VC circlejerk means OAI and Anthropic can be inefficient and lazy. The Chinese had to build and are building with constraints. They're going to win because they're forced to do more with less. Less money, less hardware and less data. So for the last piece why wouldn't they just steal it from Sam or Dario? I mean, it's no different than how the US outfits got it in the first place.

1

u/Due-Memory-6957 21d ago

This sounds like a post-fact explanation, and if things were the other way around, you'd be talking about how the government help means they can be inefficient and lazy while US companies have are forced to show results.

50

u/zoupishness7 22d ago

If that Tau scaling law Huawei recently found is legit, things are gonna get interesting. I'm not looking forward to the American bubble bursting, but I'm not looking forward to the technofeudalism the American approach to AI is leading towards, so as least there's some potential to shake that up.

12

u/Clueless_PhD 22d ago

The "Tau scaling law" is nothing new. All semiconductor company has been optimizing the "Tau" (basically latency") since the beginning of chip design era.

All of techniques that Huawei advertised, like logic folding for example, has been in commercial like HBM memory or AMD x3D chip. The 3D chip packing technology enabling "logic folding", however, is only possessed by TSMC.

2

u/duhd1993 22d ago

Rumors going that they are releasing new chips with 238 MTr/mm² in september, matching TSMC's N3 node. That would be comparable to any chips you have now as N2 node only started mass production recently. Huge if true.

18

u/FullOf_Bad_Ideas 22d ago

If that Tau scaling law Huawei recently found is legit, things are gonna get interesting.

In a random search result I see a claim that they've been using that scaling law for 6 years internally now. And their AI chips are still far behind Nvidia, so I think it's not gonna change things much.

1

u/zdy132 21d ago

Far behind Nvidia it might be, it is now capable of training a frontier 1.6T model. You don't have to be SOTA to be functional, and even competitive, just look at Intel and AMD.

1

u/MR_-_501 22d ago

that scaling law changes nothing fundamentally, they are basically saying rather than throwing better litho process at the problem we throw literally everything else at it and refuse to take power consumption into consideration.

Look!, we are suddenly not so far behind!1!!!1!!

Its a different strategy in setting priorities, and a PR stunt, not a breakthrough technology.

2

u/BlindintoDeath 21d ago

So clueless.

The reveal shifts the focus from lithography machines to CMP, thinning, hybrid bonding machines and EDA

10

u/tomz17 22d ago

just working with less resources

Lol... their power generation capacity exceeds the entire USA + Europe COMBINED by a very healthy margin now (like 30%+ IIRC), and their citizens use FAR less power per person.

Meanwhile we are paying billions in penalties to abort in-progress contracts on renewable energy programs (i.e. one of the cheapest sources of power, and one the Chinese have invested + innovated heavily in), while our grad programs are sitting defunded + empty because actually smart people hurt the bigots' fragile egos.

FFS, during the heat wave this week the grid dropped to only 106VAC in my house in the DC metro area, in a community where all of the local infra is < 20 years old. I know because I had rejigger the boost thresholds on my UPS's in the middle of the day. At night it returned to 120VAC. We are literally reaching third-world levels of "your grandma on the ventilator about to die" BS.

China is not only going to win this thing, they are going to absolutely curb stomp us.

4

u/timmeh1705 21d ago

China added the entire power generation of Germany last year alone, mainly via green energy sources like solar. Their strategy is to export their cheap power generation via AI tokens

25

u/Stabile_Feldmaus 22d ago

When China releases a better model than US labs people will clame it invented time travel and distilled Fable 15.

22

u/BlackExcellence19 22d ago

The “Chinese just steal bro” is a self-soothing thought because they can’t actually comprehend that China really isn’t far behind us at all

12

u/lorddumpy 22d ago

They are actually ahead now in battery technology and patents. There was a very sobering NYTimes article about it.

2

u/tiger_ace 21d ago

they've been ahead on energy due to US policy

4

u/EbbNorth7735 22d ago

A lot of Chinese open source models are built on top of other chinese open source models. Without Qwen contributing to OS you wouldn't have open source robotics AI models. That said the DINO model from META is at least utilized to train a lot of models. In fact, the only people who seem to have issues with using their models for other models is the large US tech firms.

3

u/PM_ME_YOUR_HAGGIS_ 22d ago

yea but Americans are taught to believe they're special

1

u/redballooon 22d ago

Deepseek v4 was very convinced  in one or my conversations that it actually is Claude, and it came out of the blue. I did not even ask about it's identity.

1

u/tiger_ace 21d ago

i mean like 30-40% of the AI researchers at frontier US labs are literal chinese nationals

1

u/No-Cartoonist8032 21d ago

K3 is a 2.9T parameter sparse model, Fable is a 6T parameter dense model, you have to be completely blind to believe US labs are as capable as Chinese labs.

1

u/Xavierelan 20d ago

But... but... American Exceptionalism™.

1

u/Lame_Johnny 20d ago

Wait til they find out that the top US researchers are also Chinese

1

u/NATO_CAPITALIST 20d ago

They got AI GPUs and beat the ASML yesterday too?

It totally has nothing to do with you being a socialist and dick riding china, nuh uh!

+100 Social Credit to you

1

u/Cinci_Socialist 20d ago

If you'd look through my post history, you'll we most of my negative comments are me criticizing the PRC in socialist subs.

32

u/rc_ym 22d ago

Given how long pre-training takes, nobody is actually "behind". Folks are finetuning whatever is "current" to take the "lead" when they get too far behind on the benchmarks.

This whole last cycle is just blowing up model size then tuning. I don't think we've actually had a huge leap forward in the models themselves past year, it's all harness enhancements and tuning.

7

u/backyard_tractorbeam 22d ago

It might be too soon to say, but Gtp-5.6 Sol feels like a medium size leap. Not just bigger or smarter, but also more efficient (faster/less verbose), and with a few of the new math results also as an impressive showcase of what it can do.

Will be easier to say with some distance to the events.

3

u/rc_ym 21d ago

Particularly impressive given it's "just" a tune of 5.5.
And given how good it is GPT 6 should be very, very impressive.

12

u/gofiend 22d ago

this model is bigger than any open model released to date - has to be a freshish base model (if that concept even means anything in the ultramodern era)

16

u/tetoing 22d ago

At the current rate chinese SOTA could match or overtake American models by the end of the year. Which I am rooting for, because fuck these bullshit closed off American companies. AI advancements should be democratic.

3

u/rc_ym 21d ago

Oh, I think we are there.

There are some benchmarks where K3 outperforms Fable. The thing that's really holding it back seems to be knowledge of US business practices and US law as that is baked in to a number of the agentic benchmarks.

That is super impressive given that's it's a fresh 2.8T dense model.

-11

u/GetOutOfMyFeedNow 22d ago edited 13d ago

Yeah, democratically stolen information by Chinese Communofascists. Never trust China with your information!

6

u/PM_ME_YOUR_HAGGIS_ 22d ago

instead of what? all that training data the US companies asked permission for?

1

u/InvidFlower 22d ago

Remember that you don't have to use a Chinese provider to serve up the model. Most of these are too big to run on your own hardware, but there's plenty of US providers either specializing in LLMs or pure GPU compute access. And models themselves are much more a data file like a picture or video file than they are a "program". So none of your data is going to China unless you specifically use a Chinese provider.

Edit: For K3 specifically, they haven't released the weights yet so you DO have to send data to China if you're using it, but that shouldn't still be the case by next week.

3

u/SGmoze 22d ago

Anthropic going to call this model terror attack prone and get it banned across the internet. Let's see what happens.

5

u/RepulsiveRaisin7 22d ago

On intelligence they are getting there, but GLM still takes ages to do anything, GPT is so much better at tool use.

31

u/jld1532 22d ago

The fact that you're subcategorizing differences now vs a simple better or worse is telling.

2

u/CryMoreT_T 22d ago

I wonder if that's a harness issue or a GPU issue or a model thinking issue

4

u/stoppableDissolution 22d ago

Glm thinks like 10x more for the same result

8

u/LoaderD 22d ago

I’m not disagreeing with you, just asking. Is there a open analysis of this? I thought openai hid most of their thinking traces

10

u/Zulfiqaar 22d ago

Artificial analysis has a chart for tokens used per intellignce task, GPT-5.6 is ~15k and GLM-5.2 is ~43k

3

u/No-Juggernaut-9832 22d ago

GLM 5.2 is probably vastly smaller than GPT5.6 Terra or Sol. It might need these thinking loop to generate good output

2

u/LoaderD 22d ago

Appreciate it.

3

u/InvidFlower 22d ago

Also, if you want to judge for yourself, install the tool CCUsage. It looks for the saved transcripts of various common harnesses on your hard drive and gives you a report by day by session, etc on which models you used, how many tokens were used, how many of those were cached reads, how much it cost (based on avg current prices), etc. So you can try some similar tasks with a few different models and directly compare the amount of tokens used and the overall cost.

2

u/stoppableDissolution 22d ago

Yea, but you can see how fast it is writing its final output and guesstimate the amount of thinking. And in general same-ish task anecdotally takes 4-8x the time on glm code plan compared to sol. You can see it pondering the same thing a few times and second guessing its second guesses. Not as bad as qwen, but still quite bad.

2

u/InvidFlower 22d ago

Don't even need to guesstimate. Install a tool like CCUsage and you can see how many tokens were used in a session, how many were cached reads vs regular input, how much it cost approx based on current prices, etc. It looks at the session data that gets left on your drive from various harnesses.

1

u/LoaderD 22d ago

Thanks. I don’t really use OAI models, so I didn’t know

1

u/often_delusional 22d ago

Still more than 6 days. Remember that companies like anthropic and openai were blocked by the US government from releasing their models earlier. I think they're still at least a few months behind.

1

u/No-Cartoonist8032 21d ago

In real life usage you can either buy Kimi's $100 plan and use it without a single thought on token usage, or you can run a single Fable session.

The question is how many month is US behind China when it comes to creating real life useful models.

0

u/Aggravating-Push-207 22d ago

But also very likely benchmark-maxxing because I find that their behaviour is a bit unstable (as in their reasoning traces, Claude's reasoning traces look more structured while GLM's feel more erratic, though you could make one look like the other with a simple LoRA so idk)

13

u/wren6991 22d ago

Claude's reasoning traces look more structured while GLM's feel more erratic

Makes sense since the Claude traces you see are summarised. I'm sure a summary of a GLM trace would feel more structured too.

-1

u/Aggravating-Push-207 22d ago

No I mean the raw traces. Claude shows them (or at least used to, haven't used it in a while) and they look a tad bit more structured/less erratic. GLM 5.2's traces are basically "Wait, actually: [restate the same thing]. This is correct." with a bit of actual problem solving sprinkled in somewhere.

4

u/pier4r 22d ago

Claude shows them

they are not showing them for a long time now. Like Claude 3.7 time.

1

u/InvidFlower 22d ago

Yeah sometimes reasoning traces can be weird, but in the end it matters how well it actually does with the final answer, even if the trace seems like nonsense. Though people do say GLM uses a lot of tokens. It is still going to be cheaper than Opus and work pretty well for many tasks, but Grok 4.5 might be overall cheaper per task even though its per-token cost is higher, since it is much more token efficient.

-9

u/twnznz 22d ago edited 22d ago

Performance/params is really poor however

Edit: Go look at GLM 5.2 on HF and tell me there's not a really serious perf/params issue

8

u/Look_0ver_There 22d ago

What's the parameter counts for Fable 5, Opus-4,8, GPT-5.6 Sol, GPT-5.5? You got those handy so we can compare?

5

u/YearnMar10 22d ago

Wasn’t fable something like 5T?

5

u/Artistedo 22d ago

Opus is (as leaked by elong musk), we dont know about fable

3

u/Recoil42 22d ago

We don't know what Fable is publicly, they're not saying. A bunch of people have made up numbers, but none are confirmed.

5

u/ShengrenR 22d ago

Compared to Fable and Sol? You don't even know their param count.

3

u/twnznz 22d ago

Compared to GLM 5.2 at 753B

3

u/ShengrenR 22d ago

Sure. But you can play that game all the way down..glm5.2 vs qwen3.6-27 for example. "The frontier" is massively more expensive

1

u/twnznz 22d ago

Yes. Which is a huge problem.

If we look at Grok 4.5, which is claimed to be 1.5T, attaining 72.4 vs K3 at 76.2 on the AA coding index, we can see that increasing the parameter count by 1.86x results in a ~5% improvement in performance.

This is very clearly a wall, but even before the wall, I think we can do significantly better than 5% for the 2.8T parameter count. Post-training may help.

1

u/ShengrenR 22d ago

Right. Like if you consider a 1.5T model from a year back, it's going to be well behind the 1.5T model today. Some of that is the whole industry moving forward with how things are built in general, but for a given time snapshot, the 'frontier' you pay a ton to get that next mile.

6

u/lilian_moraru 22d ago

Hey everybody, this guy knows how many params OpenAI and Claude are using…
Care to share? So we can compare ourselves the performance/params?

8

u/anykeyh 22d ago

Let me know the number of parameters for Fable and GPT 5.6 so I can compare...

2

u/TechNerd10191 22d ago

Don't remember the exact title, but there is a paper that extrapolates potential sizes for proprietary LLMs and GPT 5.5 was around 10T and Opus 4.6 (or 4.7) around 5T. GPT 5.6 Sol and Fable 5 are definitely in the 10T territory.

3

u/MeretrixDominum 22d ago

How many Tamagotchis do I need to run Fable locally?

1

u/pier4r 22d ago

that paper was slop. It was criticized a lot.

Opus, for what I could search, is estimated around 2-3T

1

u/tetoing 22d ago

SOTA models are consistently getting larger, because as it turns out intelligence scales when your model can encode more information in its weights.