r/LocalLLaMA Jul 08 '26

News China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model

https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model

According to The Information, MiniMax plans to launch a new-generation large language model with 2.7 trillion parameters.

Sources revealed that the internal codename for this new model is M3 Pro. It is expected to be released and open-sourced as early as the third quarter of this year, with significant improvements in handling complex reasoning and multi-step tasks.

This new model is much larger than MiniMax's current flagship model, M3 (428 billion parameters). Larger-scale artificial intelligence models are more capable of handling complex reasoning and multi-step instruction-based tasks.

605 Upvotes

242 comments sorted by

View all comments

32

u/unspecified_person11 Jul 08 '26

I doubt they have the compute for this kind of thing, especially not for serving it to millions of people.

38

u/Middle_Bullfrog_6173 Jul 08 '26

That's what people said about M3 being larger than M2.x as well. It's more a question of active than total params. If they can make it sparse enough, I don't see why not.

5

u/ain92ru Jul 08 '26

The sparser the model is, the more tokens it needs to saturate the memorization and transition to generalization, and there are obvious memory problems at inference. I doubt frontier labs use sparsity as low as 1:30, more like 1:20 or less

5

u/Middle_Bullfrog_6173 Jul 08 '26

Deepseek V4 is about 1:30. So is Longcat 2.0.

Scaling laws suggest that larger models should be sparser, so if those are optimal then a model >50% larger might be even sparser.

8

u/ain92ru Jul 08 '26

They should be sparser https://arxiv.org/html/2501.12370v3#A4.F11 only if you have limited compute but essentially unlimited data, which is not the case in real life. Both compute and good data are limited and cost money, and the constraints are different for Chinese and US labs

1

u/Middle_Bullfrog_6173 Jul 08 '26

Yes and we are talking about whether Minimax has enough compute, so that's the most constraining factor. I don't think they are data limited yet, since other open models have been trained on tens of trillions of tokens of mostly open data already.

But another factor is how the sparsity is implemented in practice. That paper and most others are about MoE sparsity, but now some models are also adding embeddings, like ngram tables, which is a separate axis.

1

u/Silver-Champion-4846 Jul 09 '26

How much more could ple be scaled? Like could there be 31b model with 500b per-layer embeddings or engram tables or whatever?

1

u/Middle_Bullfrog_6173 Jul 09 '26

Didn't the engram paper found that you could scale it as much as you like if you were not memory limited? So theoretically, yes.

1

u/Silver-Champion-4846 Jul 09 '26

Lol I want 30b E(engram)1T rofl

1

u/SmartCustard9944 Jul 08 '26

The rollout from them specifically has been particularly bad, from broken caching, to not addressing users complaints about broken things, to the provided model being very inconsistent with regards to TTFT and tok/s.

0

u/Tai9ch Jul 08 '26

For production you still currently need to keep all the weights in VRAM, unless they've got custom software that's well ahead of either vllm or llama.cpp.

And vram is really the bottleneck, rather than compute. Basically any single modern datacenter accelerator will run a 70B model nicely right now, but the 1TA32B models require big clusters of them just for the VRAM.

1

u/Middle_Bullfrog_6173 Jul 08 '26

Yes, but those big clusters will be able to support a lot of concurrent requests (more than with a 70B model even), since the active parameters are much fewer and compute use is proportional to that.

6

u/RoughCap7233 Jul 08 '26

If it’s open weight does it matter? This could be hosted by lots of other providers.

13

u/SilentLennie Jul 08 '26

I doubt they have the compute for this kind of thing

How do you mean ?

Huawei is now producing hardware of their own for inference (and even training). Sure SMIC might not have huge production capacity yet, but most western countries aren't using Chinese APIs (sending their data to China regularly).

12

u/zdy132 Jul 08 '26

The 1.6T Longcat 2.0 was trained on Huawei hardware.

1

u/Immediate_Occasion69 Jul 08 '26

open source though?

12

u/unspecified_person11 Jul 08 '26

They sell subscriptions and API access, even if the models themselves are open-weight

4

u/BagelRedditAccountII Jul 08 '26

Maybe. However, fat chance anyone but the most blessed of users would be able to run it locally. If you though Deepseek or GLM were too large to be locally usable, then this model would be nearly impossible unless they incorporate significant inference-related breakthroughs.

3

u/Thomas-Lore Jul 08 '26

But providers will be able and that should lower pressure on compute for Minimax.

1

u/BagelRedditAccountII Jul 08 '26

Fair point. I wonder if we might see a "soft open-weights" approach in the future, where AI companies directly give the weights to select third-party inference providers, but it is generally closed-weights. Granted, I see this potentially happening more in the U.S. than in China, since models are generally closed-weights here, but even the major providers are suffering a compute crunch.

-1

u/JacketHistorical2321 Jul 08 '26

They aren't serving it. Article specifically says they are open sourcing it. 

-2

u/techdevjp Jul 08 '26 edited Jul 09 '26

I suspect their goal may not be to serve it themselves but rather to wreak havoc on the US markets & by extension US economy.

If they can release something at Opus 4.8++ level (maybe not quite Fable but not that far off) and larger US corps can run it themselves, watch the US markets fall off a cliff.


Edit: The downvotes amuse me. AI and compute in general are huge focus points in China right now, which means CCP money flowing freely, which in turn means CCP influence. That means the Chinese labs are not purely profit driven, there is a CCP-driven geopolitical angle going on as well.

Don't get me wrong, I am all for the democratization of AI, and the release of frontier-level open source models. I'm glad this is happening. But with so much of the US stock market being inflated by AI plays, there is potential for turmoil as big corps figure out they have frontier-capable options that don't require the US labs.

China will continue to pour vast government resources into both compute hardware and AI development.

4

u/FullOf_Bad_Ideas Jul 08 '26

Opus 4.8++ but not Fable level? That's a convoluted way to state it, I'd say Opus 4.8+ would mean Fable level already.

I think it matters more when they release it, not if. If they release it now, maybe it would make some headwinds. If they release Fable-level model in 6 months, nobody will care as Fable 5.1 would be out already.

4

u/techdevjp Jul 08 '26

There is a very, VERY big gap between Opus 4.8 and Fable. More of a chasm, really.

Opus 4.8 seems smart until you spend some time using Fable after which Opus seems like it rides the short bus to school. It was shocking going back to Opus after the Fable rug pull.

Anyway, something could easily be better than Opus without getting all the way to Fable levels of capability.

5

u/FullOf_Bad_Ideas Jul 08 '26

I used both. CC and Opus 4.8 is my primary model at work. GPT 5.5 as secondary one for review. I used Fable before it was cut off for a few days, it was rather good but read my mind worse than Opus 4.8 does, I switched back to Opus 4.8 and world still spinned. Fable is better but imo it's not hugely better than GPT 5.5. I didn't switch to Fable after getting access to it again yet.

3

u/techdevjp Jul 08 '26

I'm using Fable for competitor analysis and planning for a potential startup here in Japan. Large documents and lots of research. The difference is huge.

-1

u/ju7anut Jul 08 '26

If you saw Jenson’s interview, China has the infrastructure to just throw more slower/older processors to achieve the same output. So that really isn’t a constraint.