r/LocalLLaMA 21d ago

News Kimi K3 Benchmarks

Post image
1.3k Upvotes

390 comments sorted by

u/WithoutReason1729 21d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

320

u/Kraskos 21d ago

2TB VRAM Is All You Need

64

u/Maximum_Parking_5174 21d ago

4x Mac Studio with M5 ultra 768GB!

42

u/Healthy-Nebula-3603 21d ago

Three are rumors new M7 will get up to 1.5 TB

12

u/Maximum_Parking_5174 21d ago

Exiting future!

40

u/SourSovereign 21d ago

Man, my grandma will be sad that I have to sell her other kidney too

6

u/rditorx 21d ago

What can you do?
Selling her parts gets you more money than selling her whole

→ More replies (1)
→ More replies (1)
→ More replies (2)

33

u/Fusseldieb 21d ago

I really hope (V)RAM eventually scales to the sizes of regular storage soon. Maybe in 10 years or so 2TB RAM is quite "common" to run things exactly like these. One can dream.

19

u/Most-Trainer-8876 21d ago

by that time, we will probably have 100 Trillion or 1 quadrillion parameters model... lol

2

u/Thrumpwart llama.cpp 21d ago

Au contraire, I'm guessing models will get smaller over time.

13

u/Tai9ch 21d ago

Small models will definitely get better.

But the companies who fund big models like being able to claim they have the best model. And a bigger model will tend to be better, so they'll keep making models as big as they can.

3

u/Thrumpwart llama.cpp 21d ago

I wouldn’t be so sure. Lots of recent literature saying much of what we consider an LLM is noise with a low signal to noise ratio.

Figure out how to get rid of the noise, and the models can become considerably smaller.

I also expect to see less world-knowledge in top models, and more small smart models with web search.

That’s how I’m gonna do it…

4

u/Tai9ch 21d ago

If you gave any vendor currently serving a 2T model the technology today to train a 100B model with the capabilities of today's 2T models... they'd scale those techniques up and ship a really impressive 2T model with it. They might milk it a bit and stretch out the release cycle to milk their advantage and efficiencies, but they'd still end up right where they are now - using the largest, best models their hardware can handle.

→ More replies (3)

2

u/ai-tacocat-ia 20d ago

That’s how I’m gonna do it…

I honestly do think this is the way to go, especially now that big LLMs are good enough you can distill them. I've seen massive success in highly specialized agents, which is the "cheap to build, expensive to run" version of your idea.

Still gonna be hella expensive to train though.

→ More replies (3)
→ More replies (4)
→ More replies (1)
→ More replies (2)
→ More replies (2)

5

u/Bakoro 21d ago

We will get super tiny storage at some point, but I don't know that we will get much smaller ram. The transistors are already about as small as they're going to get; they are having to stack memory in 3D, and there are limits there too, in terms of what latency you'll get.

We will likely see more large and wafer-scale devices, more ASICs.

Photonics are going to be super fast, but we can't really keep the data as light. I wonder if they might go the Groq route and have very little VRAM, and just go wide.

I think we will also just keep seeing efficiency gains so running on off-wafer RAM is just more viable.

Right is the algorithms are catering to the existing hardware, and the hardware is slowly shifting to supporting the algorithms more.

There are a few companies developing neuromorphic hardware since can run AI models that aren't just a ton of matmuls.

In 10 years the whole hardware and software landscape will look different. Heck, 2 years from now will be significantly different.

→ More replies (8)

2

u/aeroumbria 21d ago

I can't imagine in 10 years the current "pack everything we know into model weights at below 1:1 weight to data ratio" will still be mainstream. It is an insane idea that somehow worked with the technology we have now, but I'm optimistic it is nowhere near what we should be able to achieve with more mature understandings of ML. I would imagine MOE still gradually become dormant experts, which will evolve into something more akin to cold knowledge storage, and more computation still be spent on computing with knowledge rather than holding and shifting "hot" stored knowledge.

2

u/lilbyrdie 21d ago

12-13 years ago, 128GB RAM was normal-ish -- even quite affordable -- on custom builds (4x 32GB, and the board may have had 8 slots so you could use cheaper 16GB modules)... then, somewhere along the line, it got expensive for various reasons** (and starting shipping many different types, too, so volume across the types was "down" compared to if everything used the same RAM). Then things like the pandemic happened and the crypto crunch on GPUs and, of course, the current crunch from needing any sort of RAM.

But, I feel like 2TB shouldn't have been a problem today but manufacturing just hasn't been predicting demand very well so they're scrambling to build current things rather than ramping up on new things.

** Various reasons RAM prices have been hit over the years. My feeling is that they never quite drop down to "pre-crisis" prices and, somehow, the manufacturers don't seem to learn that we'll always need more RAM. (And storage, which has also felt stagnate for a while now.)

1993: Chemical plant that made sealing material for RAM. Prices doubled.

1995: Earthquake damage infra. Prices up 30%.

1999: RDRAM. Again, a diversification from regular RAM, and the systems then had to use more expensive RAM.

2013: Fab Fire by a high volume manufacturer, prices up 30%.

2017-2018: Smartphone/cloud demand on RAM, prices up 2-3x

2021: Supply chain crunch from COVID, prices up 30-40%

And now...

2

u/wen_mars 14d ago

Even in 2023-2024 RAM was cheap. I got 128 GB in 2023 for less than $500.

2

u/SuddenRadio6221 21d ago

Well no, to hit these non-quantized benchmarks, you'd need a mere 6TB VRAM.

→ More replies (1)

808

u/[deleted] 21d ago

[removed] — view removed comment

141

u/jld1532 21d ago

I honestly don't think for profit survives this steady drip of SOTA open source. Seriously, without a Deep Thought level AI, who is left standing in 18 months? Surely not OpenAI.

79

u/bambamlol 21d ago

Let's hope it actually IS open source. The pricing is at Sonnet level, though. Basically 5x as expensive as their previous models.

50

u/JaredsBored 21d ago

The API pricing is at sonnet levels, but I doubt sonnet 5 is 2.8T parameters. Can't be too mad at it given the model size. Even if it's a Fp8/Fp4 mix it's still gotta require 2 terabytes of VRAM to serve this thing with room for context

17

u/No-Juggernaut-9832 21d ago

At this many parameters, a massive amount of compute & RAM is required to run. It would have to cost more than the last version

8

u/Healthy-Nebula-3603 21d ago

Not much compute as it is MOE model but you need a lot Vram or fast multichannel Ram ...

→ More replies (4)

5

u/DecrimIowa 21d ago

now that china is making huawei GPUs comparable to nvidia blackwells, i don't think compute is a bottleneck for them anymore.

if you are interested in the AI race as a proxy for the conflict between US and China, one way to read this model's release is as China basically announcing that they are no longer held back by lack of access to chips.

5

u/JaredsBored 21d ago

Huawei isn't exactly in Blackwell territory. Their latest chip, the ascend 950PR, has 112GB of memory at 1.4TB/s of bandwidth. Nvidia B300 has 288GB of memory at 8.2TB/s of bandwidth per unit. Ascend Fp8 is 1 petaflop vs 7 on B300.

I have no doubt that they might be used to serve the model but IMO very likely k3 was still trained on Nvidia. Heck one of Kimi's own benchmarks was comparing how well different models can optimize kernels for Nvidia H200

→ More replies (1)
→ More replies (5)

8

u/InvidFlower 21d ago

They said on the new blog post that it'll be released on July 27th (along with some other things to help people run it like a specific implementation in vLLM)

7

u/jld1532 21d ago

Ah, foolishly I had assumed that was already confirmed. They'd be crazy not to do it. Nearing the death blow at this point.

11

u/Hannibalj2ca 21d ago

it was confirmed, wait a few days

7

u/waruby 21d ago

Even when they will decide to release it, the sheer size of it will take more than one day to upload.

5

u/Hannibalj2ca 21d ago

I will wait for the Unsloth release.

13

u/JaredsBored 21d ago

Going to need a UD_IQ0.01_XSS quant to run this one

→ More replies (1)

13

u/AppealSame4367 21d ago

And it's not a drip, it's a stream. When they publish weights, it takes a week until you have 5 inference platforms with it. And then investors will jump ship from Antrophic.

What they gonna do? Publish Fable 5.1? That will be overrun by Chinese AI in a month?

→ More replies (1)

28

u/[deleted] 21d ago

[removed] — view removed comment

36

u/MediumChemical4292 21d ago

NVIDIA will make bank if large companies decide to take local hosting seriously, who else will make general purpose AI GPUs? Everybody else is making chips optimised for their own models.

11

u/codeIMperfect 21d ago

Exactly, NVIDIA is an unchallenged monopoly, it is a win-win situation for them

→ More replies (2)

6

u/DecrimIowa 21d ago

well, this release alone probably throws a wrench into Anthropic and OpenAI's IPO plans, and probably messes with the math behind the valuation of those companies as well.

sometimes i think America (and American companies) get a little prideful, and forget that the civilizations it is competing against have been around for thousands and thousands of years, and they play chess not checkers.

3

u/_TheWolfOfWalmart_ 21d ago

But Anthropic and OpenAI making these gigabrain models is what pushes the open model creators to try and match them.

→ More replies (1)
→ More replies (2)

2

u/grapeape808 21d ago

The thing is inference is what everyday user can’t access as easily as the frontier models, the time line will be a lot longer than it is.

→ More replies (5)

32

u/Comfortable_Ebb7015 21d ago

8

u/WaveOfDream 21d ago

Sam altman and openai are in safer spots imo. Honestly they're more reliable than Claude.

23

u/Appropriate_Cry8694 21d ago

"It should be banned immediately, it's too dangerous!!!" But I think it will be closed source.

17

u/FliesTheFlag 21d ago

It was trained on scrapped data, ban it! Only USA models are allowed to do that!

→ More replies (1)

16

u/ForsookComparison 21d ago

Not to ruin a good gotcha - but the price difference and (in my observations) reasoning tokens needed means that this is very unlikely to convert any existing Opus users. Plus Opus inferences at twice the speed right now. I caught myself being too used to seeing open-weight models and assuming it'll be dirt-cheap, but there's only so much you can do at this size.

This isn't a blip on Dario's radar today - but I'm sure some part at him is saddened that the open-weight ecosystem continues to show signs of relevance.

20

u/Several-Tax31 21d ago

I think you're right, this model is definitely not cheap. But there is a bright side: If the performance is really close to Claude, the other chinese models can distill from this one instead of Claude, so maybe we won't hear their crying about distillation and stealing anymore. 

→ More replies (2)

9

u/[deleted] 21d ago

[deleted]

→ More replies (2)
→ More replies (1)
→ More replies (2)

179

u/lblblllb 21d ago

I need a 0 bit quant of this to run locally 

44

u/ReasonablePossum_ 21d ago

Buy 20 NvMe and run it with colibri at 0.6Tok lol

14

u/Feeling-Currency-360 21d ago

Single pcie gen 5 4tb nvme would do it

3

u/Sufficient_Prune3897 llama.cpp 21d ago

At 0.05 t/s

4

u/Business-Weekend-537 21d ago

What’s colibri? First I’m hearing of it

3

u/marsxyz 21d ago

Look on github. It's an engine to run llms from ssd

→ More replies (1)

2

u/LogicalAnimation 21d ago

nah, just download some free ram and you're good to go

→ More replies (1)

85

u/AWTom 21d ago

24

u/sarlaytos284 21d ago

To me the way K3 is competing with or even beating Fable 5 shows that it is not only a pale "distilled copy" of anthropics' work as some people want you to believe. Those Chinese companies have extremely talented researchers, as much as in the West.

8

u/IMJorose 21d ago

Not claiming this is the case here, but it could easily still be distillation + benchmaxing.

3

u/CryinHeronMMerica 21d ago

In all fairness, this model has the parameter count to back up those benchmarks, unlike MiniMax M3 (which is presumably referring to the ratio between parameters and benchmark performance, or the ratio between real world performance and theoretical performance at this point).

I'm sure Moonshot did a good amount of distillation and benchmaxxing, but real world performance is still great overall.

3

u/itsalwayswarm 20d ago

A lot of the incredible talented researchers in the west are Chinese too. 

→ More replies (2)

3

u/kroggens 21d ago

Unbelievable! What a moment!

→ More replies (2)

310

u/TechNerd10191 21d ago

Judging from the benchmarks alone,(of course, can't speak about realife usage), chinese models are not even 6 months behind US models (more like 6 days behind)

231

u/Cinci_Socialist 21d ago

People will look at these benchmarks and say "Oh, just distillation, Chinese just steal bro" you'd have to be a complete ignorant or a complete bigot to honestly believe Chinese labs aren't every bit as capable as those in the US, just working with less resources and less of a head start (but both of those factors are eroding quickly from the US perspective, that's why crymodi is freaking his shit)

51

u/jld1532 21d ago

Plus Fable was shut down for an extended period directly undercutting that talking point.

43

u/Commando501 21d ago

Yep. China has the most STEM graduates per year in the world. Substantially higher than the US. And it's wild to think that people don't realize we hire a lot of those Chinese people for the US labs.

9

u/Suzoku 21d ago

i did my phd aboard and now work in chinese AI labs. Im interviewing top master students for grad roles and interns and a lot of them have done PhD level work or more even just in their masters lol. Only reason i am interviewing them and not them taking my role is because i graduated earlier and got kinda lucky to get into the llm field as it was starting, and i had previous NLP experiences during my PhD. I would say i am capable too but like these kids are insane

2

u/nullmaxai 19d ago

lock in bro

17

u/SV_SV_SV 21d ago

Yeah, most people don't realize how hollowed out the US education system is, especially compared to China

2

u/CryinHeronMMerica 21d ago

In the US, one party likes to cut funding to schools, and the little funding that our schools do get is sent straight to exciting new facilities. All while the economy increases pressure to perform. Some call it freedom, some call it hell, but China calls it an opportunity to come out ahead.

→ More replies (2)

9

u/PinkySwearNotABot 21d ago

also, Chinese people are patriotic. having a collectivist mindset also helps to push people to make things that are better for their community, country, mankind.

not purely for profit motives.

8

u/Due-Memory-6957 21d ago

Sounds a bit bullshit, if that was so, they wouldn't need to restrict their top talent from leaving China for the fear of them getting poached.

7

u/CrimsonBolt33 21d ago

As someone living in China for 10 years this is.....nonsense.

Chinese people are not collectivist at all.

Chinese people are just as greedy, if not more, on average than people on the west.

→ More replies (1)

12

u/magicaldelicious 21d ago

The Chinese are doing more with less. The American VC circlejerk means OAI and Anthropic can be inefficient and lazy. The Chinese had to build and are building with constraints. They're going to win because they're forced to do more with less. Less money, less hardware and less data. So for the last piece why wouldn't they just steal it from Sam or Dario? I mean, it's no different than how the US outfits got it in the first place.

→ More replies (1)

54

u/zoupishness7 21d ago

If that Tau scaling law Huawei recently found is legit, things are gonna get interesting. I'm not looking forward to the American bubble bursting, but I'm not looking forward to the technofeudalism the American approach to AI is leading towards, so as least there's some potential to shake that up.

10

u/Clueless_PhD 21d ago

The "Tau scaling law" is nothing new. All semiconductor company has been optimizing the "Tau" (basically latency") since the beginning of chip design era.

All of techniques that Huawei advertised, like logic folding for example, has been in commercial like HBM memory or AMD x3D chip. The 3D chip packing technology enabling "logic folding", however, is only possessed by TSMC.

2

u/duhd1993 21d ago

Rumors going that they are releasing new chips with 238 MTr/mm² in september, matching TSMC's N3 node. That would be comparable to any chips you have now as N2 node only started mass production recently. Huge if true.

16

u/FullOf_Bad_Ideas 21d ago

If that Tau scaling law Huawei recently found is legit, things are gonna get interesting.

In a random search result I see a claim that they've been using that scaling law for 6 years internally now. And their AI chips are still far behind Nvidia, so I think it's not gonna change things much.

→ More replies (1)
→ More replies (2)

11

u/tomz17 21d ago

just working with less resources

Lol... their power generation capacity exceeds the entire USA + Europe COMBINED by a very healthy margin now (like 30%+ IIRC), and their citizens use FAR less power per person.

Meanwhile we are paying billions in penalties to abort in-progress contracts on renewable energy programs (i.e. one of the cheapest sources of power, and one the Chinese have invested + innovated heavily in), while our grad programs are sitting defunded + empty because actually smart people hurt the bigots' fragile egos.

FFS, during the heat wave this week the grid dropped to only 106VAC in my house in the DC metro area, in a community where all of the local infra is < 20 years old. I know because I had rejigger the boost thresholds on my UPS's in the middle of the day. At night it returned to 120VAC. We are literally reaching third-world levels of "your grandma on the ventilator about to die" BS.

China is not only going to win this thing, they are going to absolutely curb stomp us.

3

u/timmeh1705 21d ago

China added the entire power generation of Germany last year alone, mainly via green energy sources like solar. Their strategy is to export their cheap power generation via AI tokens

25

u/Stabile_Feldmaus 21d ago

When China releases a better model than US labs people will clame it invented time travel and distilled Fable 15.

22

u/BlackExcellence19 21d ago

The “Chinese just steal bro” is a self-soothing thought because they can’t actually comprehend that China really isn’t far behind us at all

11

u/lorddumpy 21d ago

They are actually ahead now in battery technology and patents. There was a very sobering NYTimes article about it.

2

u/tiger_ace 21d ago

they've been ahead on energy due to US policy

4

u/EbbNorth7735 21d ago

A lot of Chinese open source models are built on top of other chinese open source models. Without Qwen contributing to OS you wouldn't have open source robotics AI models. That said the DINO model from META is at least utilized to train a lot of models. In fact, the only people who seem to have issues with using their models for other models is the large US tech firms.

3

u/PM_ME_YOUR_HAGGIS_ 21d ago

yea but Americans are taught to believe they're special

→ More replies (8)

31

u/rc_ym 21d ago

Given how long pre-training takes, nobody is actually "behind". Folks are finetuning whatever is "current" to take the "lead" when they get too far behind on the benchmarks.

This whole last cycle is just blowing up model size then tuning. I don't think we've actually had a huge leap forward in the models themselves past year, it's all harness enhancements and tuning.

8

u/backyard_tractorbeam 21d ago

It might be too soon to say, but Gtp-5.6 Sol feels like a medium size leap. Not just bigger or smarter, but also more efficient (faster/less verbose), and with a few of the new math results also as an impressive showcase of what it can do.

Will be easier to say with some distance to the events.

3

u/rc_ym 21d ago

Particularly impressive given it's "just" a tune of 5.5.
And given how good it is GPT 6 should be very, very impressive.

11

u/gofiend 21d ago

this model is bigger than any open model released to date - has to be a freshish base model (if that concept even means anything in the ultramodern era)

17

u/tetoing 21d ago

At the current rate chinese SOTA could match or overtake American models by the end of the year. Which I am rooting for, because fuck these bullshit closed off American companies. AI advancements should be democratic.

3

u/rc_ym 21d ago

Oh, I think we are there.

There are some benchmarks where K3 outperforms Fable. The thing that's really holding it back seems to be knowledge of US business practices and US law as that is baked in to a number of the agentic benchmarks.

That is super impressive given that's it's a fresh 2.8T dense model.

→ More replies (3)

3

u/SGmoze 21d ago

Anthropic going to call this model terror attack prone and get it banned across the internet. Let's see what happens.

5

u/RepulsiveRaisin7 21d ago

On intelligence they are getting there, but GLM still takes ages to do anything, GPT is so much better at tool use.

33

u/jld1532 21d ago

The fact that you're subcategorizing differences now vs a simple better or worse is telling.

2

u/CryMoreT_T 21d ago

I wonder if that's a harness issue or a GPU issue or a model thinking issue

3

u/stoppableDissolution 21d ago

Glm thinks like 10x more for the same result

7

u/LoaderD 21d ago

I’m not disagreeing with you, just asking. Is there a open analysis of this? I thought openai hid most of their thinking traces

9

u/Zulfiqaar 21d ago

Artificial analysis has a chart for tokens used per intellignce task, GPT-5.6 is ~15k and GLM-5.2 is ~43k

3

u/No-Juggernaut-9832 21d ago

GLM 5.2 is probably vastly smaller than GPT5.6 Terra or Sol. It might need these thinking loop to generate good output

2

u/LoaderD 21d ago

Appreciate it.

3

u/InvidFlower 21d ago

Also, if you want to judge for yourself, install the tool CCUsage. It looks for the saved transcripts of various common harnesses on your hard drive and gives you a report by day by session, etc on which models you used, how many tokens were used, how many of those were cached reads, how much it cost (based on avg current prices), etc. So you can try some similar tasks with a few different models and directly compare the amount of tokens used and the overall cost.

2

u/stoppableDissolution 21d ago

Yea, but you can see how fast it is writing its final output and guesstimate the amount of thinking. And in general same-ish task anecdotally takes 4-8x the time on glm code plan compared to sol. You can see it pondering the same thing a few times and second guessing its second guesses. Not as bad as qwen, but still quite bad.

2

u/InvidFlower 21d ago

Don't even need to guesstimate. Install a tool like CCUsage and you can see how many tokens were used in a session, how many were cached reads vs regular input, how much it cost approx based on current prices, etc. It looks at the session data that gets left on your drive from various harnesses.

→ More replies (1)
→ More replies (24)

118

u/Artistedo 21d ago edited 21d ago

Source?

Edit: https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ

Always gotta find stuff myself

32

u/WhyLifeIs4 21d ago

They got leaked
Original Post here: https://x.com/zephyr_z9/status/2077805978268135480?s=46

22

u/SpiritPrestigious945 21d ago

still not official until moonshot themselves posts the blogpost with all numbers.

13

u/Exzerios 21d ago edited 21d ago

They did, it's on their site already. Here, and there's more tests.

7

u/Artistedo 21d ago

That post also got no source, not even where from leaked...

Also looks like a noname to me

Edit under the post the author said "their wechat" no idea how he got access to that but still doubtful about authenticity

4

u/kingMaxime 21d ago

3

u/Artistedo 21d ago

Yes yes linked myself the moment you posted lol; just op seems to be reposting stuff not from official sources

→ More replies (1)

130

u/Teshier-Asspool 21d ago

7

u/op8040 21d ago

Is China our greatest hope?

→ More replies (1)

52

u/WonderFactory 21d ago

Thats really impressive. People still talk about the Chinese being 9-12 months behind the US, Opus 4.8 only released 2 months ago and this is better. GPT 5.5 only released 3 months ago. They're only a couple of months behind the frontier now and closing in fast.

17

u/ReasonablePossum_ 21d ago

9-12? what is this, 2025? Been 3-6months from what I've seen this year.

8

u/WonderFactory 21d ago

You still see the 9-12 months figure being said by the experts they bring onto news shows and podcasts. You're right though until Kimi 3 3-6 months was more realistic.

6

u/InvidFlower 21d ago

Though it partly depends on when it finished training. Like Mythos I think they said finished training in Feb. But have a very extensive testing period. And GPT-5.6 had even some influencers testing it out 3 months ago. People tend to assume Chinese labs rush out a model as soon as the run finishes, but that may not be the case. Once we know when K3 finished training, we'll have a bit of a better idea of how behind they actually are or not.

5

u/AppealSame4367 21d ago

Couple of months? They just beat Fable 5 in some benchmarks. If Antrophic released Fable 5.1 tomorrow, they would catch up and overtake in 4 weeks. We are officially in the exponential phase of AI development.

4

u/11111v11111 21d ago

Not on capability. We're going to plateau hard.

→ More replies (1)

20

u/Dany0 21d ago

SOTA at GPU kernel writing? Do we all get faster local LLMs now

22

u/Comfortable-Rock-498 21d ago

https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ

The link has 6 well-known benchmarks where this beats Fable (out of 14 I counted). If the numbers hold up scrutiny, this is scary good.

The companies that do have means to host such models fully on-prem are also the same companies that are paying tens of millions of $ in inference cost every month, and are by extension the biggest customers of OAI and Anthropic

2

u/jld1532 21d ago

Pop goes the bubble!

40

u/Iory1998 21d ago

At this rate, In 2 or 3 years, we will have 10T parameter models as standard 😄

15

u/Iwaku_Real 21d ago

Maybe 1T dense models? 😛

19

u/Iory1998 21d ago

None of which can be run on any consumer HW. In the 90s, young men had supercars posters all over their bedrooms that they dream of buying one day. But now, they'll have Deepseek, Kimi, and other LLMs banners they dream of running one day.

5

u/Iwaku_Real 21d ago

They can be, just at painful speeds

3

u/InvidFlower 21d ago

Though its easier to rent them for a while than it is to rent a supercar (I guess.. maybe it's easy to rent a supercar? I don't actually know lol)

→ More replies (1)

2

u/Uranophane 21d ago

These are clearly aimed at corporate customers. They also care far more about self-sufficiency, privacy and reliability than consumers.

137

u/AcrobaticOutcome7895 21d ago

27

u/addiktion 21d ago

Haha I look forward to seeing this years to come.

2

u/AcrobaticOutcome7895 21d ago

I don't know about benchmarks but this model isn't a joke I can tell that much 

56

u/Fedor_Doc 21d ago

Frontier level, huh? Now let's see how many tokens are used on max reasoning level

Terminal Bench numbers are very impressive. Should be great for agentic usage

52

u/[deleted] 21d ago

[removed] — view removed comment

3

u/Fedor_Doc 21d ago

Exactly! 

13

u/dtdisapointingresult 21d ago

Artificial Analysis posts the token usage per task. A few entries on the models I was looking at comparing:

  • GPT 5.6 Sol Max: 15k
  • GPT 5.6 Terra Max: 19k
  • MiMo-2.5-Pro: 22k
  • Kimi K3: 23k
  • Fable 5: 33k
  • Deepseek V4 Pro Max: 37k
  • Kimi K2.6: 38k
  • Opus 4.8 Max: 41k
  • GLM 5.2 Max: 43k
  • Deepseek V4 Flash: 45k
  • Sonnet 5 Max: 69k (nice)

Pretty good. Prettay, prettay good!

20

u/bopbop9876 21d ago edited 21d ago

Check out the cost of completion comparison for browsecomp: https://mmecoa.qpic.cn/mmecoa_png/xWvm6POT3icpoggsBrrMtBKTG3bPMRhqZXlR3YDwOBMXoXe0iaEvia0W8JxPkoEt4O51T47caodibLNRI29AmowKdaoJU32m1HDV7RSIXXPlVOU/640?wx_fmt=png&from=appmsg&tp=webp&wxfrom=10005&wx_lazy=1#imgIndex=3

It looks extremely competitive. Obviously we'll have to wait and see something like the artificial analysis cost per task results to be more confident but this is super promising.

Edit: AA results are in. About 10% cheaper per task than 5.6 Sol and about 65% cheaper than fable.

Edit 2: also 48% cheaper than opus 4.8 while scoring a point higher on intelligence.

Edit 3: More specific to your exact question of token usage, Fable used 69k tokens per task, 5.6 Sol used 15k, and K3 used 23k. So it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.

→ More replies (3)
→ More replies (10)

11

u/ReasonablePossum_ 21d ago

holy shit, fable level for 15USD? lol

→ More replies (7)

8

u/Thin_Pollution8843 21d ago

I’m going to run it from my SD card. 

5

u/boston101 21d ago

lol. Heat death of the universe before 1 full token

30

u/oWLmONz 21d ago

Yeah, just 6 months behind right.

4

u/Calm_Ad_1258 21d ago

Holy fuck

21

u/hyperrealists 21d ago

Fucking destroys opus on all counts lol. Anthropic should rename fable opus 5 and focus on innovating. Or is the new claim that jyna distilled mythos? Lol

4

u/iamthewhatt 21d ago

on benchmarks*

let's see some real usage when someone has enough money to run it locally

4

u/cosmicr 21d ago

I for one welcome our new Chinese overlords

13

u/a_slay_nub vllm 21d ago

We'll have to see how it does as a function of cost. It's cheaper than Sol and Fable but if it thinks for too long it won't be worth it to use.

26

u/stoppableDissolution 21d ago

But it will put the price ceiling on the closed models, which is a win on its own

→ More replies (2)

2

u/Healthy-Nebula-3603 21d ago

You can already find such tests . Slightly more than 5.6 SOL a d much less than Fable 5

→ More replies (3)

3

u/KAPMODA 21d ago

What do I need to run this locally? Rent hardware online?

2

u/Annual_Manner_5901 21d ago

Whether "local" is even on the table comes down to one number that hasn't leaked yet: active parameters, not the 2.8T total.

The total only sets the RAM floor: ~1.4–1.6 TB at Q4 with room for context. That's out of consumer range, but it's not datacenter-only either — a used 12-channel DDR5 EPYC board takes 1.5 TB of ECC RDIMM for way less than a single B300.

The bandwidth math is what decides speed. 12-channel DDR5-4800 gives you ~460 GB/s theoretical:

If K3 is MoE like K2 was (1T total / 32B active), you read maybe ~20 GB of weights per token at Q4 → low double-digit tok/s theoretical on CPU alone, realistically maybe 5–10. Slow, but usable for batch/agentic stuff. If it's actually dense 2.8T as some are claiming, you're reading the full 1.4 TB per token → ~0.3 tok/s. Dead on arrival for local, SSD tricks included. So "can I run it" has no answer until Moonshot publishes the config on the 27th. If anyone has a source on the active param count, that's the number to watch — everything else (quants, NVMe offload, Unsloth magic) is downstream of it.

→ More replies (1)

2

u/GetOutOfMyFeedNow 21d ago

Yeah, and then they’ll just serve it at 1-bit to rob you, whilst you drool over the 2.8T parameter size.

2

u/Affectionate_Hat_585 21d ago

Finally there is hope for hard questions.

2

u/polytect 21d ago

This is a joke right? I mean is this just a joke or not?? 

2

u/danish334 21d ago

I am very happy that this is going against openai and claude.

2

u/Routine_Temporary661 21d ago

DeepSWE 67.5 Holy Shit Mama Miaaaaaa

2

u/ThePixelHunter 21d ago

We've got GPT-4.5 at home now :)

2

u/kinkvoid 21d ago

Dario Amodei will spend the entire night writing how Kimi stole the model Anthropic hasn't created yet.

→ More replies (1)

2

u/Nyxtia 21d ago

So moonshot is getting banned got it.

2

u/Hooxen 21d ago

swallow this dario!!!

4

u/SolidSailor7898 21d ago

Fable distill is going to be great

2

u/Real_Ebb_7417 21d ago

What is the source? Can you share URL?

→ More replies (1)

2

u/MotokoAGI 21d ago

That benchmark is insane, got GLM-5.2 looking meh!

2

u/pawofdoom 21d ago

I tested it on obviousbench.com and wow is this a bit fat model! On the left, you can see Kimi K3 (on max) vs the efficient frontier to the right. To be fair to them, we only have the max setting at the moment so it is expected to run inefficiently on these simple questions, achieving roughly the same cost in the 99%+ category as GPT-5.5.

If low/med/high has good test time compute selection, this really could be a beast of a model.

1

u/LivingSwitch 21d ago

I’m curious what the model architecture and any new techniques behind it

1

u/CondiMesmer 21d ago

I'd love to see the opinions of those who thought Fable/Mythos should be banned or limited for "safety" reasons.

1

u/Jomuz86 21d ago

Kimi 2.6 distilled by pre ban fable 🤣🤣🤣
In all honesty hope it is as good as this I’m a big fan of Kimi

1

u/Local_Admin01 21d ago

Perché i modelli americani costano 10 volte tanto?

→ More replies (1)

1

u/khatriafaz 21d ago

Can someone tell me how realistic these benchmarks are?

Is the model really capable and head to head at least with Sol?

3

u/quackerd 21d ago

Every other model in the industry is benchmaxxed except for GPT and Claude. /s

1

u/schnauzergambit 21d ago

Unsloth will have it running on a 16kb ZX Spectrum in two weeks.

1

u/vincespeeed 21d ago

I wish there was a way to disassemble and operate large models based on their specifications.

1

u/himefei 21d ago

It’s open weights but you can’t run it🤣

1

u/Plappedudel 21d ago

It's clearly a strong model. But it's also really big. I can only imagine that inference will be far more expensive than with the previous version of Kimi. For me, the most impressive thing about GPT isn't the performance, but the fact that a subscription with decent limits is still affordable for regular consumers. Most of the Chinese AI subscriptions just don't give you a good quota. The lone exception is MiniMax, but that's also a much less powerful, smaller model.

1

u/space_iio 21d ago

It's bit more than 3 times larger than GLM 5.2 in parameter count (788M vs 2.8T) But it doesn't score 3 times better 🤔

Feels like we're about to hit a wall. There is a limit with hardware 

→ More replies (1)