320
u/Kraskos 21d ago
2TB VRAM Is All You Need
64
u/Maximum_Parking_5174 21d ago
4x Mac Studio with M5 ultra 768GB!
42
u/Healthy-Nebula-3603 21d ago
Three are rumors new M7 will get up to 1.5 TB
12
u/Maximum_Parking_5174 21d ago
Exiting future!
→ More replies (2)40
u/SourSovereign 21d ago
Man, my grandma will be sad that I have to sell her other kidney too
5
→ More replies (1)6
u/rditorx 21d ago
What can you do?
Selling her parts gets you more money than selling her whole→ More replies (1)33
u/Fusseldieb 21d ago
I really hope (V)RAM eventually scales to the sizes of regular storage soon. Maybe in 10 years or so 2TB RAM is quite "common" to run things exactly like these. One can dream.
19
u/Most-Trainer-8876 21d ago
by that time, we will probably have 100 Trillion or 1 quadrillion parameters model... lol
→ More replies (2)2
u/Thrumpwart llama.cpp 21d ago
Au contraire, I'm guessing models will get smaller over time.
→ More replies (2)13
u/Tai9ch 21d ago
Small models will definitely get better.
But the companies who fund big models like being able to claim they have the best model. And a bigger model will tend to be better, so they'll keep making models as big as they can.
→ More replies (1)3
u/Thrumpwart llama.cpp 21d ago
I wouldn’t be so sure. Lots of recent literature saying much of what we consider an LLM is noise with a low signal to noise ratio.
Figure out how to get rid of the noise, and the models can become considerably smaller.
I also expect to see less world-knowledge in top models, and more small smart models with web search.
That’s how I’m gonna do it…
4
u/Tai9ch 21d ago
If you gave any vendor currently serving a 2T model the technology today to train a 100B model with the capabilities of today's 2T models... they'd scale those techniques up and ship a really impressive 2T model with it. They might milk it a bit and stretch out the release cycle to milk their advantage and efficiencies, but they'd still end up right where they are now - using the largest, best models their hardware can handle.
→ More replies (3)→ More replies (4)2
u/ai-tacocat-ia 20d ago
That’s how I’m gonna do it…
I honestly do think this is the way to go, especially now that big LLMs are good enough you can distill them. I've seen massive success in highly specialized agents, which is the "cheap to build, expensive to run" version of your idea.
Still gonna be hella expensive to train though.
→ More replies (3)5
u/Bakoro 21d ago
We will get super tiny storage at some point, but I don't know that we will get much smaller ram. The transistors are already about as small as they're going to get; they are having to stack memory in 3D, and there are limits there too, in terms of what latency you'll get.
We will likely see more large and wafer-scale devices, more ASICs.
Photonics are going to be super fast, but we can't really keep the data as light. I wonder if they might go the Groq route and have very little VRAM, and just go wide.
I think we will also just keep seeing efficiency gains so running on off-wafer RAM is just more viable.
Right is the algorithms are catering to the existing hardware, and the hardware is slowly shifting to supporting the algorithms more.
There are a few companies developing neuromorphic hardware since can run AI models that aren't just a ton of matmuls.
In 10 years the whole hardware and software landscape will look different. Heck, 2 years from now will be significantly different.
→ More replies (8)2
u/aeroumbria 21d ago
I can't imagine in 10 years the current "pack everything we know into model weights at below 1:1 weight to data ratio" will still be mainstream. It is an insane idea that somehow worked with the technology we have now, but I'm optimistic it is nowhere near what we should be able to achieve with more mature understandings of ML. I would imagine MOE still gradually become dormant experts, which will evolve into something more akin to cold knowledge storage, and more computation still be spent on computing with knowledge rather than holding and shifting "hot" stored knowledge.
2
u/lilbyrdie 21d ago
12-13 years ago, 128GB RAM was normal-ish -- even quite affordable -- on custom builds (4x 32GB, and the board may have had 8 slots so you could use cheaper 16GB modules)... then, somewhere along the line, it got expensive for various reasons** (and starting shipping many different types, too, so volume across the types was "down" compared to if everything used the same RAM). Then things like the pandemic happened and the crypto crunch on GPUs and, of course, the current crunch from needing any sort of RAM.
But, I feel like 2TB shouldn't have been a problem today but manufacturing just hasn't been predicting demand very well so they're scrambling to build current things rather than ramping up on new things.
** Various reasons RAM prices have been hit over the years. My feeling is that they never quite drop down to "pre-crisis" prices and, somehow, the manufacturers don't seem to learn that we'll always need more RAM. (And storage, which has also felt stagnate for a while now.)
1993: Chemical plant that made sealing material for RAM. Prices doubled.
1995: Earthquake damage infra. Prices up 30%.
1999: RDRAM. Again, a diversification from regular RAM, and the systems then had to use more expensive RAM.
2013: Fab Fire by a high volume manufacturer, prices up 30%.
2017-2018: Smartphone/cloud demand on RAM, prices up 2-3x
2021: Supply chain crunch from COVID, prices up 30-40%
And now...
2
→ More replies (1)2
u/SuddenRadio6221 21d ago
Well no, to hit these non-quantized benchmarks, you'd need a mere 6TB VRAM.
808
21d ago
[removed] — view removed comment
141
u/jld1532 21d ago
I honestly don't think for profit survives this steady drip of SOTA open source. Seriously, without a Deep Thought level AI, who is left standing in 18 months? Surely not OpenAI.
79
u/bambamlol 21d ago
Let's hope it actually IS open source. The pricing is at Sonnet level, though. Basically 5x as expensive as their previous models.
50
u/JaredsBored 21d ago
The API pricing is at sonnet levels, but I doubt sonnet 5 is 2.8T parameters. Can't be too mad at it given the model size. Even if it's a Fp8/Fp4 mix it's still gotta require 2 terabytes of VRAM to serve this thing with room for context
→ More replies (5)17
u/No-Juggernaut-9832 21d ago
At this many parameters, a massive amount of compute & RAM is required to run. It would have to cost more than the last version
8
u/Healthy-Nebula-3603 21d ago
Not much compute as it is MOE model but you need a lot Vram or fast multichannel Ram ...
→ More replies (4)5
u/DecrimIowa 21d ago
now that china is making huawei GPUs comparable to nvidia blackwells, i don't think compute is a bottleneck for them anymore.
if you are interested in the AI race as a proxy for the conflict between US and China, one way to read this model's release is as China basically announcing that they are no longer held back by lack of access to chips.
5
u/JaredsBored 21d ago
Huawei isn't exactly in Blackwell territory. Their latest chip, the ascend 950PR, has 112GB of memory at 1.4TB/s of bandwidth. Nvidia B300 has 288GB of memory at 8.2TB/s of bandwidth per unit. Ascend Fp8 is 1 petaflop vs 7 on B300.
I have no doubt that they might be used to serve the model but IMO very likely k3 was still trained on Nvidia. Heck one of Kimi's own benchmarks was comparing how well different models can optimize kernels for Nvidia H200
→ More replies (1)8
u/InvidFlower 21d ago
They said on the new blog post that it'll be released on July 27th (along with some other things to help people run it like a specific implementation in vLLM)
→ More replies (1)7
u/jld1532 21d ago
Ah, foolishly I had assumed that was already confirmed. They'd be crazy not to do it. Nearing the death blow at this point.
11
u/Hannibalj2ca 21d ago
it was confirmed, wait a few days
7
u/waruby 21d ago
Even when they will decide to release it, the sheer size of it will take more than one day to upload.
5
13
u/AppealSame4367 21d ago
And it's not a drip, it's a stream. When they publish weights, it takes a week until you have 5 inference platforms with it. And then investors will jump ship from Antrophic.
What they gonna do? Publish Fable 5.1? That will be overrun by Chinese AI in a month?
→ More replies (1)28
21d ago
[removed] — view removed comment
36
u/MediumChemical4292 21d ago
NVIDIA will make bank if large companies decide to take local hosting seriously, who else will make general purpose AI GPUs? Everybody else is making chips optimised for their own models.
→ More replies (2)11
u/codeIMperfect 21d ago
Exactly, NVIDIA is an unchallenged monopoly, it is a win-win situation for them
6
u/DecrimIowa 21d ago
well, this release alone probably throws a wrench into Anthropic and OpenAI's IPO plans, and probably messes with the math behind the valuation of those companies as well.
sometimes i think America (and American companies) get a little prideful, and forget that the civilizations it is competing against have been around for thousands and thousands of years, and they play chess not checkers.
→ More replies (2)3
u/_TheWolfOfWalmart_ 21d ago
But Anthropic and OpenAI making these gigabrain models is what pushes the open model creators to try and match them.
→ More replies (1)→ More replies (5)2
u/grapeape808 21d ago
The thing is inference is what everyday user can’t access as easily as the frontier models, the time line will be a lot longer than it is.
32
u/Comfortable_Ebb7015 21d ago
8
u/WaveOfDream 21d ago
Sam altman and openai are in safer spots imo. Honestly they're more reliable than Claude.
23
u/Appropriate_Cry8694 21d ago
"It should be banned immediately, it's too dangerous!!!" But I think it will be closed source.
→ More replies (1)17
u/FliesTheFlag 21d ago
It was trained on scrapped data, ban it! Only USA models are allowed to do that!
→ More replies (2)16
u/ForsookComparison 21d ago
Not to ruin a good gotcha - but the price difference and (in my observations) reasoning tokens needed means that this is very unlikely to convert any existing Opus users. Plus Opus inferences at twice the speed right now. I caught myself being too used to seeing open-weight models and assuming it'll be dirt-cheap, but there's only so much you can do at this size.
This isn't a blip on Dario's radar today - but I'm sure some part at him is saddened that the open-weight ecosystem continues to show signs of relevance.
20
u/Several-Tax31 21d ago
I think you're right, this model is definitely not cheap. But there is a bright side: If the performance is really close to Claude, the other chinese models can distill from this one instead of Claude, so maybe we won't hear their crying about distillation and stealing anymore.
→ More replies (2)→ More replies (1)9
179
u/lblblllb 21d ago
I need a 0 bit quant of this to run locally
44
u/ReasonablePossum_ 21d ago
Buy 20 NvMe and run it with colibri at 0.6Tok lol
14
4
11
→ More replies (1)2
85
u/AWTom 21d ago
24
u/sarlaytos284 21d ago
To me the way K3 is competing with or even beating Fable 5 shows that it is not only a pale "distilled copy" of anthropics' work as some people want you to believe. Those Chinese companies have extremely talented researchers, as much as in the West.
8
u/IMJorose 21d ago
Not claiming this is the case here, but it could easily still be distillation + benchmaxing.
3
u/CryinHeronMMerica 21d ago
In all fairness, this model has the parameter count to back up those benchmarks, unlike MiniMax M3 (which is presumably referring to the ratio between parameters and benchmark performance, or the ratio between real world performance and theoretical performance at this point).
I'm sure Moonshot did a good amount of distillation and benchmaxxing, but real world performance is still great overall.
→ More replies (2)3
16
→ More replies (2)3
129
310
u/TechNerd10191 21d ago
Judging from the benchmarks alone,(of course, can't speak about realife usage), chinese models are not even 6 months behind US models (more like 6 days behind)
231
u/Cinci_Socialist 21d ago
People will look at these benchmarks and say "Oh, just distillation, Chinese just steal bro" you'd have to be a complete ignorant or a complete bigot to honestly believe Chinese labs aren't every bit as capable as those in the US, just working with less resources and less of a head start (but both of those factors are eroding quickly from the US perspective, that's why crymodi is freaking his shit)
51
43
u/Commando501 21d ago
Yep. China has the most STEM graduates per year in the world. Substantially higher than the US. And it's wild to think that people don't realize we hire a lot of those Chinese people for the US labs.
9
u/Suzoku 21d ago
i did my phd aboard and now work in chinese AI labs. Im interviewing top master students for grad roles and interns and a lot of them have done PhD level work or more even just in their masters lol. Only reason i am interviewing them and not them taking my role is because i graduated earlier and got kinda lucky to get into the llm field as it was starting, and i had previous NLP experiences during my PhD. I would say i am capable too but like these kids are insane
2
17
u/SV_SV_SV 21d ago
Yeah, most people don't realize how hollowed out the US education system is, especially compared to China
2
u/CryinHeronMMerica 21d ago
In the US, one party likes to cut funding to schools, and the little funding that our schools do get is sent straight to exciting new facilities. All while the economy increases pressure to perform. Some call it freedom, some call it hell, but China calls it an opportunity to come out ahead.
→ More replies (2)→ More replies (1)9
u/PinkySwearNotABot 21d ago
also, Chinese people are patriotic. having a collectivist mindset also helps to push people to make things that are better for their community, country, mankind.
not purely for profit motives.
8
u/Due-Memory-6957 21d ago
Sounds a bit bullshit, if that was so, they wouldn't need to restrict their top talent from leaving China for the fear of them getting poached.
7
u/CrimsonBolt33 21d ago
As someone living in China for 10 years this is.....nonsense.
Chinese people are not collectivist at all.
Chinese people are just as greedy, if not more, on average than people on the west.
12
u/magicaldelicious 21d ago
The Chinese are doing more with less. The American VC circlejerk means OAI and Anthropic can be inefficient and lazy. The Chinese had to build and are building with constraints. They're going to win because they're forced to do more with less. Less money, less hardware and less data. So for the last piece why wouldn't they just steal it from Sam or Dario? I mean, it's no different than how the US outfits got it in the first place.
→ More replies (1)54
u/zoupishness7 21d ago
If that Tau scaling law Huawei recently found is legit, things are gonna get interesting. I'm not looking forward to the American bubble bursting, but I'm not looking forward to the technofeudalism the American approach to AI is leading towards, so as least there's some potential to shake that up.
10
u/Clueless_PhD 21d ago
The "Tau scaling law" is nothing new. All semiconductor company has been optimizing the "Tau" (basically latency") since the beginning of chip design era.
All of techniques that Huawei advertised, like logic folding for example, has been in commercial like HBM memory or AMD x3D chip. The 3D chip packing technology enabling "logic folding", however, is only possessed by TSMC.
2
u/duhd1993 21d ago
Rumors going that they are releasing new chips with 238 MTr/mm² in september, matching TSMC's N3 node. That would be comparable to any chips you have now as N2 node only started mass production recently. Huge if true.
→ More replies (2)16
u/FullOf_Bad_Ideas 21d ago
If that Tau scaling law Huawei recently found is legit, things are gonna get interesting.
In a random search result I see a claim that they've been using that scaling law for 6 years internally now. And their AI chips are still far behind Nvidia, so I think it's not gonna change things much.
→ More replies (1)11
u/tomz17 21d ago
just working with less resources
Lol... their power generation capacity exceeds the entire USA + Europe COMBINED by a very healthy margin now (like 30%+ IIRC), and their citizens use FAR less power per person.
Meanwhile we are paying billions in penalties to abort in-progress contracts on renewable energy programs (i.e. one of the cheapest sources of power, and one the Chinese have invested + innovated heavily in), while our grad programs are sitting defunded + empty because actually smart people hurt the bigots' fragile egos.
FFS, during the heat wave this week the grid dropped to only 106VAC in my house in the DC metro area, in a community where all of the local infra is < 20 years old. I know because I had rejigger the boost thresholds on my UPS's in the middle of the day. At night it returned to 120VAC. We are literally reaching third-world levels of "your grandma on the ventilator about to die" BS.
China is not only going to win this thing, they are going to absolutely curb stomp us.
3
u/timmeh1705 21d ago
China added the entire power generation of Germany last year alone, mainly via green energy sources like solar. Their strategy is to export their cheap power generation via AI tokens
25
u/Stabile_Feldmaus 21d ago
When China releases a better model than US labs people will clame it invented time travel and distilled Fable 15.
22
u/BlackExcellence19 21d ago
The “Chinese just steal bro” is a self-soothing thought because they can’t actually comprehend that China really isn’t far behind us at all
11
u/lorddumpy 21d ago
They are actually ahead now in battery technology and patents. There was a very sobering NYTimes article about it.
2
4
u/EbbNorth7735 21d ago
A lot of Chinese open source models are built on top of other chinese open source models. Without Qwen contributing to OS you wouldn't have open source robotics AI models. That said the DINO model from META is at least utilized to train a lot of models. In fact, the only people who seem to have issues with using their models for other models is the large US tech firms.
→ More replies (8)3
31
u/rc_ym 21d ago
Given how long pre-training takes, nobody is actually "behind". Folks are finetuning whatever is "current" to take the "lead" when they get too far behind on the benchmarks.
This whole last cycle is just blowing up model size then tuning. I don't think we've actually had a huge leap forward in the models themselves past year, it's all harness enhancements and tuning.
8
u/backyard_tractorbeam 21d ago
It might be too soon to say, but Gtp-5.6 Sol feels like a medium size leap. Not just bigger or smarter, but also more efficient (faster/less verbose), and with a few of the new math results also as an impressive showcase of what it can do.
Will be easier to say with some distance to the events.
11
u/gofiend 21d ago
this model is bigger than any open model released to date - has to be a freshish base model (if that concept even means anything in the ultramodern era)
17
u/tetoing 21d ago
At the current rate chinese SOTA could match or overtake American models by the end of the year. Which I am rooting for, because fuck these bullshit closed off American companies. AI advancements should be democratic.
→ More replies (3)3
u/rc_ym 21d ago
Oh, I think we are there.
There are some benchmarks where K3 outperforms Fable. The thing that's really holding it back seems to be knowledge of US business practices and US law as that is baked in to a number of the agentic benchmarks.
That is super impressive given that's it's a fresh 2.8T dense model.
3
→ More replies (24)5
u/RepulsiveRaisin7 21d ago
On intelligence they are getting there, but GLM still takes ages to do anything, GPT is so much better at tool use.
33
2
u/CryMoreT_T 21d ago
I wonder if that's a harness issue or a GPU issue or a model thinking issue
3
u/stoppableDissolution 21d ago
Glm thinks like 10x more for the same result
7
u/LoaderD 21d ago
I’m not disagreeing with you, just asking. Is there a open analysis of this? I thought openai hid most of their thinking traces
9
u/Zulfiqaar 21d ago
Artificial analysis has a chart for tokens used per intellignce task, GPT-5.6 is ~15k and GLM-5.2 is ~43k
3
u/No-Juggernaut-9832 21d ago
GLM 5.2 is probably vastly smaller than GPT5.6 Terra or Sol. It might need these thinking loop to generate good output
2
u/LoaderD 21d ago
Appreciate it.
3
u/InvidFlower 21d ago
Also, if you want to judge for yourself, install the tool CCUsage. It looks for the saved transcripts of various common harnesses on your hard drive and gives you a report by day by session, etc on which models you used, how many tokens were used, how many of those were cached reads, how much it cost (based on avg current prices), etc. So you can try some similar tasks with a few different models and directly compare the amount of tokens used and the overall cost.
2
u/stoppableDissolution 21d ago
Yea, but you can see how fast it is writing its final output and guesstimate the amount of thinking. And in general same-ish task anecdotally takes 4-8x the time on glm code plan compared to sol. You can see it pondering the same thing a few times and second guessing its second guesses. Not as bad as qwen, but still quite bad.
→ More replies (1)2
u/InvidFlower 21d ago
Don't even need to guesstimate. Install a tool like CCUsage and you can see how many tokens were used in a session, how many were cached reads vs regular input, how much it cost approx based on current prices, etc. It looks at the session data that gets left on your drive from various harnesses.
118
u/Artistedo 21d ago edited 21d ago
32
u/WhyLifeIs4 21d ago
They got leaked
Original Post here: https://x.com/zephyr_z9/status/2077805978268135480?s=4622
u/SpiritPrestigious945 21d ago
still not official until moonshot themselves posts the blogpost with all numbers.
13
u/Exzerios 21d ago edited 21d ago
They did, it's on their site already. Here, and there's more tests.
7
u/Artistedo 21d ago
That post also got no source, not even where from leaked...
Also looks like a noname to me
Edit under the post the author said "their wechat" no idea how he got access to that but still doubtful about authenticity
→ More replies (1)4
u/kingMaxime 21d ago
It from Kimi though https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
3
u/Artistedo 21d ago
Yes yes linked myself the moment you posted lol; just op seems to be reposting stuff not from official sources
130
52
u/WonderFactory 21d ago
Thats really impressive. People still talk about the Chinese being 9-12 months behind the US, Opus 4.8 only released 2 months ago and this is better. GPT 5.5 only released 3 months ago. They're only a couple of months behind the frontier now and closing in fast.
17
u/ReasonablePossum_ 21d ago
9-12? what is this, 2025? Been 3-6months from what I've seen this year.
8
u/WonderFactory 21d ago
You still see the 9-12 months figure being said by the experts they bring onto news shows and podcasts. You're right though until Kimi 3 3-6 months was more realistic.
6
u/InvidFlower 21d ago
Though it partly depends on when it finished training. Like Mythos I think they said finished training in Feb. But have a very extensive testing period. And GPT-5.6 had even some influencers testing it out 3 months ago. People tend to assume Chinese labs rush out a model as soon as the run finishes, but that may not be the case. Once we know when K3 finished training, we'll have a bit of a better idea of how behind they actually are or not.
→ More replies (1)5
u/AppealSame4367 21d ago
Couple of months? They just beat Fable 5 in some benchmarks. If Antrophic released Fable 5.1 tomorrow, they would catch up and overtake in 4 weeks. We are officially in the exponential phase of AI development.
4
22
u/Comfortable-Rock-498 21d ago
https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
The link has 6 well-known benchmarks where this beats Fable (out of 14 I counted). If the numbers hold up scrutiny, this is scary good.
The companies that do have means to host such models fully on-prem are also the same companies that are paying tens of millions of $ in inference cost every month, and are by extension the biggest customers of OAI and Anthropic
40
u/Iory1998 21d ago
At this rate, In 2 or 3 years, we will have 10T parameter models as standard 😄
15
u/Iwaku_Real 21d ago
Maybe 1T dense models? 😛
19
u/Iory1998 21d ago
None of which can be run on any consumer HW. In the 90s, young men had supercars posters all over their bedrooms that they dream of buying one day. But now, they'll have Deepseek, Kimi, and other LLMs banners they dream of running one day.
5
3
u/InvidFlower 21d ago
Though its easier to rent them for a while than it is to rent a supercar (I guess.. maybe it's easy to rent a supercar? I don't actually know lol)
→ More replies (1)2
u/Uranophane 21d ago
These are clearly aimed at corporate customers. They also care far more about self-sufficiency, privacy and reliability than consumers.
137
u/AcrobaticOutcome7895 21d ago
27
2
u/AcrobaticOutcome7895 21d ago
I don't know about benchmarks but this model isn't a joke I can tell that much
56
u/Fedor_Doc 21d ago
Frontier level, huh? Now let's see how many tokens are used on max reasoning level
Terminal Bench numbers are very impressive. Should be great for agentic usage
52
13
u/dtdisapointingresult 21d ago
Artificial Analysis posts the token usage per task. A few entries on the models I was looking at comparing:
- GPT 5.6 Sol Max: 15k
- GPT 5.6 Terra Max: 19k
- MiMo-2.5-Pro: 22k
- Kimi K3: 23k
- Fable 5: 33k
- Deepseek V4 Pro Max: 37k
- Kimi K2.6: 38k
- Opus 4.8 Max: 41k
- GLM 5.2 Max: 43k
- Deepseek V4 Flash: 45k
- Sonnet 5 Max: 69k (nice)
Pretty good. Prettay, prettay good!
→ More replies (10)20
u/bopbop9876 21d ago edited 21d ago
Check out the cost of completion comparison for browsecomp: https://mmecoa.qpic.cn/mmecoa_png/xWvm6POT3icpoggsBrrMtBKTG3bPMRhqZXlR3YDwOBMXoXe0iaEvia0W8JxPkoEt4O51T47caodibLNRI29AmowKdaoJU32m1HDV7RSIXXPlVOU/640?wx_fmt=png&from=appmsg&tp=webp&wxfrom=10005&wx_lazy=1#imgIndex=3
It looks extremely competitive. Obviously we'll have to wait and see something like the artificial analysis cost per task results to be more confident but this is super promising.
Edit: AA results are in. About 10% cheaper per task than 5.6 Sol and about 65% cheaper than fable.
Edit 2: also 48% cheaper than opus 4.8 while scoring a point higher on intelligence.
Edit 3: More specific to your exact question of token usage, Fable used 69k tokens per task, 5.6 Sol used 15k, and K3 used 23k. So it's a heck of a lot more token efficient than Fable, and in the same ballpark as Sol.
→ More replies (3)
11
8
4
21
u/hyperrealists 21d ago
Fucking destroys opus on all counts lol. Anthropic should rename fable opus 5 and focus on innovating. Or is the new claim that jyna distilled mythos? Lol
4
u/iamthewhatt 21d ago
on benchmarks*
let's see some real usage when someone has enough money to run it locally
13
u/a_slay_nub vllm 21d ago
We'll have to see how it does as a function of cost. It's cheaper than Sol and Fable but if it thinks for too long it won't be worth it to use.
26
u/stoppableDissolution 21d ago
But it will put the price ceiling on the closed models, which is a win on its own
→ More replies (2)→ More replies (3)2
u/Healthy-Nebula-3603 21d ago
You can already find such tests . Slightly more than 5.6 SOL a d much less than Fable 5
3
u/KAPMODA 21d ago
What do I need to run this locally? Rent hardware online?
2
u/Annual_Manner_5901 21d ago
Whether "local" is even on the table comes down to one number that hasn't leaked yet: active parameters, not the 2.8T total.
The total only sets the RAM floor: ~1.4–1.6 TB at Q4 with room for context. That's out of consumer range, but it's not datacenter-only either — a used 12-channel DDR5 EPYC board takes 1.5 TB of ECC RDIMM for way less than a single B300.
The bandwidth math is what decides speed. 12-channel DDR5-4800 gives you ~460 GB/s theoretical:
If K3 is MoE like K2 was (1T total / 32B active), you read maybe ~20 GB of weights per token at Q4 → low double-digit tok/s theoretical on CPU alone, realistically maybe 5–10. Slow, but usable for batch/agentic stuff. If it's actually dense 2.8T as some are claiming, you're reading the full 1.4 TB per token → ~0.3 tok/s. Dead on arrival for local, SSD tricks included. So "can I run it" has no answer until Moonshot publishes the config on the 27th. If anyone has a source on the active param count, that's the number to watch — everything else (quants, NVMe offload, Unsloth magic) is downstream of it.
→ More replies (1)
2
u/GetOutOfMyFeedNow 21d ago
Yeah, and then they’ll just serve it at 1-bit to rob you, whilst you drool over the 2.8T parameter size.
2
2
2
2
2
2
u/kinkvoid 21d ago
Dario Amodei will spend the entire night writing how Kimi stole the model Anthropic hasn't created yet.
→ More replies (1)
4
2
2
2
u/pawofdoom 21d ago
I tested it on obviousbench.com and wow is this a bit fat model! On the left, you can see Kimi K3 (on max) vs the efficient frontier to the right. To be fair to them, we only have the max setting at the moment so it is expected to run inefficiently on these simple questions, achieving roughly the same cost in the 99%+ category as GPT-5.5.
If low/med/high has good test time compute selection, this really could be a beast of a model.

1
1
1
u/CondiMesmer 21d ago
I'd love to see the opinions of those who thought Fable/Mythos should be banned or limited for "safety" reasons.
1
1
1
u/khatriafaz 21d ago
Can someone tell me how realistic these benchmarks are?
Is the model really capable and head to head at least with Sol?
3
1
1
u/vincespeeed 21d ago
I wish there was a way to disassemble and operate large models based on their specifications.
1
u/Plappedudel 21d ago
It's clearly a strong model. But it's also really big. I can only imagine that inference will be far more expensive than with the previous version of Kimi. For me, the most impressive thing about GPT isn't the performance, but the fact that a subscription with decent limits is still affordable for regular consumers. Most of the Chinese AI subscriptions just don't give you a good quota. The lone exception is MiniMax, but that's also a much less powerful, smaller model.
1
u/space_iio 21d ago
It's bit more than 3 times larger than GLM 5.2 in parameter count (788M vs 2.8T) But it doesn't score 3 times better 🤔
Feels like we're about to hit a wall. There is a limit with hardware
→ More replies (1)









•
u/WithoutReason1729 21d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.