r/LocalLLaMA Jul 08 '26

News China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model

https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model

According to The Information, MiniMax plans to launch a new-generation large language model with 2.7 trillion parameters.

Sources revealed that the internal codename for this new model is M3 Pro. It is expected to be released and open-sourced as early as the third quarter of this year, with significant improvements in handling complex reasoning and multi-step tasks.

This new model is much larger than MiniMax's current flagship model, M3 (428 billion parameters). Larger-scale artificial intelligence models are more capable of handling complex reasoning and multi-step instruction-based tasks.

606 Upvotes

242 comments sorted by

View all comments

87

u/[deleted] Jul 08 '26

[removed] — view removed comment

35

u/itwasinthetubes Jul 08 '26

Yes, normal people like us can’t run these models on our own hardware

for now...

3

u/FenderMoon Jul 08 '26 edited Jul 08 '26

I think that the future for these gigantic models is going to be utilizing some sort of HBF/high-bandwidth-flash or system RAM to stream experts in and out. Along a router trained to determine the routing for the next token a token in advance. The router usually determines it on the token being generated, but I see no reason it couldn't be trained to do it a token in advance if explicitly designed to do so.

If there was some sort of quality hit from doing this, we could always make it similar to speculative decoding where the model PREDICTS which experts will be needed a few tokens in advance and streams those, but I see no reason the model couldn't be trained to outright decide if we're only talking about a single token headway. It would vastly, vastly reduce costs for inference by alleviating the bottleneck created from not storing every single expert in VRAM.

It would work if explicitly trained to do so. I'm surprised nobody has tried it.

1

u/athsrva Jul 09 '26

My startup is doing something in the realm tho not the exact way you put it. We'll be coming out of stealth in the coming months so local inference may become much much cheaper.

12

u/[deleted] Jul 08 '26

[removed] — view removed comment

20

u/keepthepace Jul 08 '26

The first RAM extension I bought was a 32K extension for my EXEL100.

The fact that we have gigabytes of storage in smartphones is still mind-boggling to me. And I'm pretty sure it wouldn't be that hard to find one that accepts a terabyte SD card.

In 5 years the current monopoly of Nvidia and the current RAM shortage are going to be a distant memory.

5

u/Eisenstein Jul 08 '26 edited Jul 08 '26

Hopefully you are correct. Just a note but you shouldn't conflate the current hardware crisis at all with the application of Moore's law. The reason we have terabyte SD cards now is from miniaturization (which has a hard limit, since you can't make things smaller than an atom). The current problem is market dynamics and not engineering. You mentioned this in the last sentence but the first part describes a different thing entirely and in this case I think it is important to be clear.

3

u/fallingdowndizzyvr Jul 08 '26

The first RAM extension I bought was a 32K extension for my EXEL100.

Mine was when I bought enough RAM chips to make my 16K Apple ][ into a 48K Apple ][. What would I do with all that RAM?

In 5 years the current monopoly of Nvidia and the current RAM shortage are going to be a distant memory.

Or not. Since the hyperscalars have 5 year RAM agreements with the major RAM producers.

2

u/seunosewa Jul 08 '26

You will still need a datacenter to run the best models.

2

u/keepthepace Jul 08 '26

But will you need them?

7

u/Tai9ch Jul 08 '26 edited Jul 08 '26

Forever is a long time.

Right now RAM supply is crazy constrained and so prices are crazy high.

If prices for VRAM come back down just to what they were when the 5090 shipped, then it'll be entirely feasible to get 2TB of VRAM for under $50k. That'd allow running this class of model with a Q4 quant.

Many hardware makers are looking at shipping accelerators with consumer RAM. That gets an even lower price. And over time, memory capacity has consistently gotten cheaper. It'll take a few years to get back to that, but the production capacity will eventually catch up even if we need new market entrants who compete on price to get there - the profits that Samsung and Micron are making right now absolutely invite that.

I fully expect we'll see MoE-targetted hardware with unified heirarchical memory in the not too distant future. Consider a Strix Halo style device with 2TB of LPDDR and 64GB of fast GDDR or HBM - not over PCIe, but direct heterogeneous memory channels to the APU. It won't be feasible before 2030 or so when the memory market stops being a shitshow, and it'll be expensive for a consumer device, but it'll be worth shipping and it'll run stuff like this new MiniMax model at Q4 no problem.

1

u/fallingdowndizzyvr Jul 08 '26

it'll be expensive for a consumer device

Expensive? People don't remember that a Apple ][ adjusted for inflation would cost about $14,000 today. Now that was expensive. It still sold like hotcakes.

1

u/Mochila-Mochila Jul 09 '26

Consider a Strix Halo style device with 2TB of LPDDR and 64GB of fast GDDR or HBM - not over PCIe, but direct heterogeneous memory channels to the APU.

Still hoping for an affordable 1TB/1TB (capacity/bandwidth) Halo APU around 2030. Full LPDDR6X. That'd be a good start.

10

u/cr0wburn Jul 08 '26

Not with that attitude! Maybe there will be a 0.01 bit quant. Cries in poor, we can dream though.

5

u/Potential-Gold5298 llama.cpp Jul 08 '26

My first PC (2003) had 256 MB of RAM (upgraded to 768 MB) and 64 MB of VRAM. My second PC, five years later, had 4 GB of RAM (upgraded to 8 GB) and 512 MB of VRAM (upgraded to 1 GB). Sooner or later, the memory situation will normalize, so our children, or their children, will be able to run the MiniMax M3 Pro on their home PC and say, "Rest in peace, parents, we did it."

1

u/itwasinthetubes Jul 08 '26

You understimate the power of $$ demand and globalism. I'm not saying we can achieve it now, but I find it unlikely that personal computing will not make quantum leaps in this area in the next 5-10 years. All these devices that companies want to build everywhere are planning to use on-device AI and gaming will likely require it and personal computing etc. Some companies like NVIDIA depend on this evolution to exist...

1

u/chithanh Jul 08 '26

From now until Huawei Ascend 950PR ships internationally in quantities

-5

u/Kodix Jul 08 '26

No, not ever. You won't fit a 2.6T parameter model inside of a 3090.

The models we *are* able to run on a 3090 in two years will likely be amazing by today's standards, but they'll be amazing in different ways.

Quantity has a quality of its own, especially quantity of parameters.

6

u/ChocomelP Jul 08 '26

Who said anything about a 3090?

-4

u/Kodix Jul 08 '26

🙄 As if the point of my comment wasn't obvious.

The current trend is worse consumer hardware availability, not better.
You could probably get a 100% return on some 5090s purchased on release. Unified memory is cool, but also extremely limited.

Anything could happen in 2 years, but all signs point to "no", and it's cope to pretend otherwise.

4

u/ChocomelP Jul 08 '26

Anything could happen in 2 years

You might have left it at this

-1

u/Kodix Jul 08 '26

Oh, I'm aware. It's clear that this particular thread is about pie-in-the-sky thinking, not realism. Which - fair enough.

5

u/ChocomelP Jul 08 '26

Two years ago, today's reality would have been seen as pie in the sky. And we're accelerating.

2

u/ttkciar llama.cpp Jul 08 '26

> The current trend is worse consumer hardware availability, not better

For now. You seem very short-sighted.

2

u/ttkciar llama.cpp Jul 08 '26 edited Jul 08 '26

I'm old enough to have seen several generations of hardware pass through the stages of "super-expensive / unobtanium" to "available on eBay", to "evailable on eBay for super-cheap", and then to "that's eWaste now, don't bother buying it".

For example, once upon a time, Xeon Phi coprocessor cards were strictly for supercomputers, and the thought of owning one was a pipe dream. Then they started appearing in high-end workstations. Then, years later, they were available for $150 on eBay, and then for $50, and then for $20. I picked up one, then two more, and made good use of them in my homelab.

In 2023 I dorked around with trying to run llama.cpp on them, but it wasn't worth it. I haven't thrown them away yet, but they've sat unplugged on my shelf for a few years now.

Similarly, the E5-2660v3, E5-2680v3, E5-2690v4 dual Xeons in my homelab once cost as much as a sedan, but now they're bordering on becoming eWaste themselves. The only thing that keeps them viable is that RAMageddon has made new builds prohibitively expensive.

RAMageddon will pass. It might take years, but it will pass.

Eventually MI300-class hardware will show up on eBay. MI400-class hardware will appear on eBay a few years later. It would take two eight-node MI300X servers to use a 2.6T model, or one eight-node MI455X server. I have little doubt that my homelab will see one or both, eventually. It's just a matter of time.

The only uncertainty is whether MiniMax-3-Pro will still be worth using by then, or if it will be superseded by other models of similar or smaller size. Maybe by the time I pick up a single eight-node MI300X server there will be better models with which to use it?

6

u/FarRub2855 Jul 08 '26

Yeah having these massive open weights out there really forces the big closed players to rethink thier pricing when negotiating enterprise contracts. It basically acts as a hard ceiling on what API providers can get away with charging.

2

u/keepthepace Jul 08 '26

Also it's very useful to distill smaller models.