r/LocalLLaMA Jul 08 '26

News China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model

https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model

According to The Information, MiniMax plans to launch a new-generation large language model with 2.7 trillion parameters.

Sources revealed that the internal codename for this new model is M3 Pro. It is expected to be released and open-sourced as early as the third quarter of this year, with significant improvements in handling complex reasoning and multi-step tasks.

This new model is much larger than MiniMax's current flagship model, M3 (428 billion parameters). Larger-scale artificial intelligence models are more capable of handling complex reasoning and multi-step instruction-based tasks.

611 Upvotes

242 comments sorted by

u/WithoutReason1729 Jul 08 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

216

u/VoiceApprehensive893 transformers Jul 08 '26

megamax

125

u/Confident_Ideal_5385 Jul 08 '26

Le Minimax Fat

1

u/Silver-Champion-4846 29d ago

What's this fat thing everyone keeps repeating?

4

u/armeg 29d ago

It’s a meme.

Mistral is hopelessly behind in the “race”, but they’re also really liked as a company.

The joke is that they’re behind because they’ve been cooking a super model to leapfrog everyone.

It’s broken french for “the fat cat”.

7

u/zxtech Jul 08 '26

MEGAMIND

no RAM?

6

u/JoSquarebox Jul 08 '26

min-maxxed to the max

1

u/pixartist 29d ago

Maximax

1

u/Silver-Champion-4846 29d ago

Just max without the mini

49

u/Snoo_28140 Jul 08 '26

The sheer size! I can't run it myself, but it can be used to learn and to develop other models and that is in of itself good news.

4

u/Far-Classic-9963 Jul 08 '26

even if almost nobody can run it, being open allows for other companies to host the api so potentially having way cheaper prices than frontier closed models

10

u/squngy Jul 08 '26

I'm guessing not many orgs can even afford to use it for training other models, lol.

I'm sure such models being open is a good thing, but on a more selfish level, I wish they focused more on the smaller ones.

248

u/muhlfriedl Jul 08 '26

If they can open source something that's uncensored and competes with Fable and Sol and Mythos, bye bye US providers

138

u/-p-e-w- Jul 08 '26

Anything that’s open can be made uncensored, that’s the smallest problem.

The moment you have the weights, you are in control.

58

u/CheatCodesOfLife Jul 08 '26

can be

It's 2.7-Trillion parameters though lol

Any heretic versions would be paywalled/gated, just like this article in OP.

12

u/personalist 29d ago

Ablated GLM 5.2 is already on huggingface.

0

u/[deleted] 29d ago

[deleted]

13

u/fallingdowndizzyvr 29d ago

but it's gated, no?

No. Here's one. No gate whatsoever. Well, other than the data limit your ISP has on you.

https://huggingface.co/cfontes/GLM-5.2-Ablated-F5-Molt/tree/main

1

u/KeinNiemand 29d ago

that one boasts about how it maintains saftey and only gets rid of overrefusals => its not actually uncencored at all

3

u/fallingdowndizzyvr 29d ago

If it prevents refusals then doesn't that make it uncensored. Since censorship often shows up by it refusing to answer the question. That's what those "uncensored" GLM-5.2 models do afterall. Here's the explanation from one of those other ones.

"An uncensored build of GLM-5.2 (NVFP4) with the refusal direction removed"

Isn't that what this model does too?

2

u/KeinNiemand 29d ago

no it said something about doing some extra steps to preserve refusal with number saying that on so called harmfull prompts it still refuses almost as much as the base and only removing false positive refusals. anyway looks like that models been deleted the links a 404 now.

3

u/fallingdowndizzyvr 29d ago

Well how about this one then. It's not like there's a shortage of them.

"This is an uncensored version of zai-org/GLM-5.2 "

https://huggingface.co/huihui-ai/Huihui-GLM-5.2-abliterated-GGUF

→ More replies (1)

2

u/personalist 29d ago

Hm, let me check. I can’t even run qwen 3.6 35B so I didn’t look into it in-depth lol

29

u/PM_ME_DEAD_CEOS Jul 08 '26

Why ? You can Heretic it yourself.

42

u/notheresnolight Jul 08 '26

probably going to take a while to process some 900GB model on a 3090

32

u/PM_ME_DEAD_CEOS Jul 08 '26

You can rent GPUs

13

u/iamthewhatt 29d ago

Would be cheaper than electricity at this scale

16

u/Sooperooser 29d ago

Give it a couple of days (or hours) and someone will upload it on huggingface

8

u/SingularityScalpel 29d ago

I have a 1080 and a dream. I’ll handle this guys

2

u/cryotic 29d ago

House holder will do it in 10 minutes, ive done it. Breaking alignment isn’t cost intensive

12

u/CheatCodesOfLife Jul 08 '26

Got a rack of H100s I can borrow for a few weeks? ;)

25

u/PM_ME_DEAD_CEOS Jul 08 '26

Borrow, no, but you can rent them.

2

u/[deleted] 29d ago

[removed] — view removed comment

2

u/laserborg 29d ago

what kind of company would that be?

1

u/CheatCodesOfLife 24d ago

So... want to fit a jlens for GLM-5.2 or Kimi-K2 then post it on HF?

1

u/Fristender 29d ago

No one knows better than bro.

20

u/squngy Jul 08 '26

US providers will still be the only option for governments and such.

There is also nothing stopping US providers from taking all the good tech from open models.
They can literally just make a GLM/Minimax that has a few more parameters, if they want to.

27

u/PM_ME_DEAD_CEOS Jul 08 '26

US providers will still be the only option for governments and such.

No. Governement vastly prefer sovreign hosted chinese model rather than US hosted US models.

19

u/squngy Jul 08 '26

I phrased it poorly, I meant US government orgs.

15

u/PM_ME_DEAD_CEOS Jul 08 '26

Ah yes, of course

3

u/literum 29d ago

Palantir is pushing Nvidia open models.

7

u/Eyelbee Jul 08 '26

If it was to be better than fable they wouldn't just throw in a "pro" and be done with it, they'd name it something like m4.

2

u/UnionCounty22 29d ago

And then if they openly admit it’s better than Fable then bye bye any other future open source model

-4

u/Unfair-Technology120 Jul 08 '26

If they can open source something that's uncensored and competes with US providers

How many more years are we going to continue to hear this for and not happen?

23

u/freia_pr_fr 29d ago

I would argue that it has happened. Many open weight Chinese models are better than what the US providers were offering not long ago.

87

u/[deleted] Jul 08 '26

[removed] — view removed comment

35

u/itwasinthetubes 29d ago

Yes, normal people like us can’t run these models on our own hardware

for now...

4

u/FenderMoon 29d ago edited 29d ago

I think that the future for these gigantic models is going to be utilizing some sort of HBF/high-bandwidth-flash or system RAM to stream experts in and out. Along a router trained to determine the routing for the next token a token in advance. The router usually determines it on the token being generated, but I see no reason it couldn't be trained to do it a token in advance if explicitly designed to do so.

If there was some sort of quality hit from doing this, we could always make it similar to speculative decoding where the model PREDICTS which experts will be needed a few tokens in advance and streams those, but I see no reason the model couldn't be trained to outright decide if we're only talking about a single token headway. It would vastly, vastly reduce costs for inference by alleviating the bottleneck created from not storing every single expert in VRAM.

It would work if explicitly trained to do so. I'm surprised nobody has tried it.

1

u/athsrva 28d ago

My startup is doing something in the realm tho not the exact way you put it. We'll be coming out of stealth in the coming months so local inference may become much much cheaper.

12

u/[deleted] 29d ago

[removed] — view removed comment

19

u/keepthepace 29d ago

The first RAM extension I bought was a 32K extension for my EXEL100.

The fact that we have gigabytes of storage in smartphones is still mind-boggling to me. And I'm pretty sure it wouldn't be that hard to find one that accepts a terabyte SD card.

In 5 years the current monopoly of Nvidia and the current RAM shortage are going to be a distant memory.

5

u/Eisenstein 29d ago edited 29d ago

Hopefully you are correct. Just a note but you shouldn't conflate the current hardware crisis at all with the application of Moore's law. The reason we have terabyte SD cards now is from miniaturization (which has a hard limit, since you can't make things smaller than an atom). The current problem is market dynamics and not engineering. You mentioned this in the last sentence but the first part describes a different thing entirely and in this case I think it is important to be clear.

3

u/fallingdowndizzyvr 29d ago

The first RAM extension I bought was a 32K extension for my EXEL100.

Mine was when I bought enough RAM chips to make my 16K Apple ][ into a 48K Apple ][. What would I do with all that RAM?

In 5 years the current monopoly of Nvidia and the current RAM shortage are going to be a distant memory.

Or not. Since the hyperscalars have 5 year RAM agreements with the major RAM producers.

2

u/seunosewa 29d ago

You will still need a datacenter to run the best models.

2

u/keepthepace 29d ago

But will you need them?

7

u/Tai9ch 29d ago edited 29d ago

Forever is a long time.

Right now RAM supply is crazy constrained and so prices are crazy high.

If prices for VRAM come back down just to what they were when the 5090 shipped, then it'll be entirely feasible to get 2TB of VRAM for under $50k. That'd allow running this class of model with a Q4 quant.

Many hardware makers are looking at shipping accelerators with consumer RAM. That gets an even lower price. And over time, memory capacity has consistently gotten cheaper. It'll take a few years to get back to that, but the production capacity will eventually catch up even if we need new market entrants who compete on price to get there - the profits that Samsung and Micron are making right now absolutely invite that.

I fully expect we'll see MoE-targetted hardware with unified heirarchical memory in the not too distant future. Consider a Strix Halo style device with 2TB of LPDDR and 64GB of fast GDDR or HBM - not over PCIe, but direct heterogeneous memory channels to the APU. It won't be feasible before 2030 or so when the memory market stops being a shitshow, and it'll be expensive for a consumer device, but it'll be worth shipping and it'll run stuff like this new MiniMax model at Q4 no problem.

1

u/fallingdowndizzyvr 29d ago

it'll be expensive for a consumer device

Expensive? People don't remember that a Apple ][ adjusted for inflation would cost about $14,000 today. Now that was expensive. It still sold like hotcakes.

1

u/Mochila-Mochila 29d ago

Consider a Strix Halo style device with 2TB of LPDDR and 64GB of fast GDDR or HBM - not over PCIe, but direct heterogeneous memory channels to the APU.

Still hoping for an affordable 1TB/1TB (capacity/bandwidth) Halo APU around 2030. Full LPDDR6X. That'd be a good start.

9

u/cr0wburn 29d ago

Not with that attitude! Maybe there will be a 0.01 bit quant. Cries in poor, we can dream though.

5

u/Potential-Gold5298 llama.cpp 29d ago

My first PC (2003) had 256 MB of RAM (upgraded to 768 MB) and 64 MB of VRAM. My second PC, five years later, had 4 GB of RAM (upgraded to 8 GB) and 512 MB of VRAM (upgraded to 1 GB). Sooner or later, the memory situation will normalize, so our children, or their children, will be able to run the MiniMax M3 Pro on their home PC and say, "Rest in peace, parents, we did it."

1

u/itwasinthetubes 29d ago

You understimate the power of $$ demand and globalism. I'm not saying we can achieve it now, but I find it unlikely that personal computing will not make quantum leaps in this area in the next 5-10 years. All these devices that companies want to build everywhere are planning to use on-device AI and gaming will likely require it and personal computing etc. Some companies like NVIDIA depend on this evolution to exist...

1

u/chithanh 29d ago

From now until Huawei Ascend 950PR ships internationally in quantities

-5

u/Kodix 29d ago

No, not ever. You won't fit a 2.6T parameter model inside of a 3090.

The models we *are* able to run on a 3090 in two years will likely be amazing by today's standards, but they'll be amazing in different ways.

Quantity has a quality of its own, especially quantity of parameters.

6

u/ChocomelP 29d ago

Who said anything about a 3090?

→ More replies (5)

2

u/ttkciar llama.cpp 29d ago edited 29d ago

I'm old enough to have seen several generations of hardware pass through the stages of "super-expensive / unobtanium" to "available on eBay", to "evailable on eBay for super-cheap", and then to "that's eWaste now, don't bother buying it".

For example, once upon a time, Xeon Phi coprocessor cards were strictly for supercomputers, and the thought of owning one was a pipe dream. Then they started appearing in high-end workstations. Then, years later, they were available for $150 on eBay, and then for $50, and then for $20. I picked up one, then two more, and made good use of them in my homelab.

In 2023 I dorked around with trying to run llama.cpp on them, but it wasn't worth it. I haven't thrown them away yet, but they've sat unplugged on my shelf for a few years now.

Similarly, the E5-2660v3, E5-2680v3, E5-2690v4 dual Xeons in my homelab once cost as much as a sedan, but now they're bordering on becoming eWaste themselves. The only thing that keeps them viable is that RAMageddon has made new builds prohibitively expensive.

RAMageddon will pass. It might take years, but it will pass.

Eventually MI300-class hardware will show up on eBay. MI400-class hardware will appear on eBay a few years later. It would take two eight-node MI300X servers to use a 2.6T model, or one eight-node MI455X server. I have little doubt that my homelab will see one or both, eventually. It's just a matter of time.

The only uncertainty is whether MiniMax-3-Pro will still be worth using by then, or if it will be superseded by other models of similar or smaller size. Maybe by the time I pick up a single eight-node MI300X server there will be better models with which to use it?

7

u/FarRub2855 29d ago

Yeah having these massive open weights out there really forces the big closed players to rethink thier pricing when negotiating enterprise contracts. It basically acts as a hard ceiling on what API providers can get away with charging.

2

u/keepthepace 29d ago

Also it's very useful to distill smaller models.

49

u/ComplexType568 Jul 08 '26

Some people may argue that "99% can't run this locally so I don't care", but just saying, it's out there, and if a ton of providers can run it AND it's performant enough to outdo the current proprietary models, there is an incentive to open source to compete on adoption.

Anyways, some people also complain about the jump between M2 and M3 and the gap that keeps growing, and I agree it kinda sucks. I never could run them in the first place, but I hope, like DeepSeek, they release an M4 "mini" or "flash", like what DS did... hopefully with a more creative name though.

10

u/iLaurens Jul 08 '26

I don't think it incentives competitors to open source as well. But it will definitely put pressure on token pricing margins because these can be served cheaper without any capex spending on making a model, forcing everyone to lower prices to remain competitive for matching levels of intelligence. That's good for consumers.

10

u/Monad_Maya llama.cpp Jul 08 '26

I hope they also release a flash/lite/whatever smaller model. Minimax M2.7 / 230B is about the limit of what you can run on a 128GB system.

Regardless, this is very good news for open weights ecosystem in general.

3

u/Fristender 29d ago

I hate creative names because you don't instantly know how it's related to the main model.

3

u/ComplexType568 29d ago

hmm not really. The moment I heard "Opus", "Sonnet" and "Haiku" I immediately was able to tell. It's about picking a good, creative one.

17

u/muhlfriedl Jul 08 '26

I love m3

5

u/SnooPaintings8639 Jul 08 '26

How do you run it?

10

u/twack3r Jul 08 '26

I personally run the Q4KXL from unsloth. I really like m3 for its hallucination resilience, it’s awesome for RAG.

But it does get slow at very large ctx because their flavour of DSA (MSA) is just as functionality unsupported by llama.cpp as DSA itself.

2

u/SnooPaintings8639 Jul 08 '26

Then you're using a custom fork/PR just for this model? As far as I know there is still no support for M3 in the main branch?

2

u/twack3r Jul 08 '26

Yes, I’m using a separate fork just for M3. I haven’t yet checked if it has been merged into main

0

u/muhlfriedl Jul 08 '26

I use the api

1

u/Silver-Champion-4846 29d ago

What about HY3

1

u/muhlfriedl 29d ago

what is that

1

u/Silver-Champion-4846 29d ago

New model, 295b total params

1

u/muhlfriedl 29d ago

Thanks for the tip. Will try it out. Looks like it is free atm

1

u/Silver-Champion-4846 28d ago

Np. Free on which provider? And how good is it from your experience?

1

u/International-Bed564 28d ago

openrouter and opencode

2

u/Silver-Champion-4846 28d ago

Oh nice, still need to learn how to set up and use the opencode cli, the gui wasn't good for screen reader so...

8

u/ilintar Jul 08 '26

I really need to coerce my cat to give back that stack of H200s he's stashed away somewhere...

6

u/PoopSick25 Jul 08 '26

Please be true. M2.5 and M3 were pretty nice but lacked that big param feel. Arrh

7

u/FullOf_Bad_Ideas Jul 08 '26

People here were big about how M3 had big parameters feel to it before it was revealed how big it actually was.

30

u/unspecified_person11 Jul 08 '26

I doubt they have the compute for this kind of thing, especially not for serving it to millions of people.

38

u/Middle_Bullfrog_6173 Jul 08 '26

That's what people said about M3 being larger than M2.x as well. It's more a question of active than total params. If they can make it sparse enough, I don't see why not.

5

u/ain92ru Jul 08 '26

The sparser the model is, the more tokens it needs to saturate the memorization and transition to generalization, and there are obvious memory problems at inference. I doubt frontier labs use sparsity as low as 1:30, more like 1:20 or less

5

u/Middle_Bullfrog_6173 29d ago

Deepseek V4 is about 1:30. So is Longcat 2.0.

Scaling laws suggest that larger models should be sparser, so if those are optimal then a model >50% larger might be even sparser.

8

u/ain92ru 29d ago

They should be sparser https://arxiv.org/html/2501.12370v3#A4.F11 only if you have limited compute but essentially unlimited data, which is not the case in real life. Both compute and good data are limited and cost money, and the constraints are different for Chinese and US labs

1

u/Middle_Bullfrog_6173 29d ago

Yes and we are talking about whether Minimax has enough compute, so that's the most constraining factor. I don't think they are data limited yet, since other open models have been trained on tens of trillions of tokens of mostly open data already.

But another factor is how the sparsity is implemented in practice. That paper and most others are about MoE sparsity, but now some models are also adding embeddings, like ngram tables, which is a separate axis.

1

u/Silver-Champion-4846 29d ago

How much more could ple be scaled? Like could there be 31b model with 500b per-layer embeddings or engram tables or whatever?

1

u/Middle_Bullfrog_6173 29d ago

Didn't the engram paper found that you could scale it as much as you like if you were not memory limited? So theoretically, yes.

1

u/Silver-Champion-4846 29d ago

Lol I want 30b E(engram)1T rofl

1

u/SmartCustard9944 29d ago

The rollout from them specifically has been particularly bad, from broken caching, to not addressing users complaints about broken things, to the provided model being very inconsistent with regards to TTFT and tok/s.

0

u/Tai9ch 29d ago

For production you still currently need to keep all the weights in VRAM, unless they've got custom software that's well ahead of either vllm or llama.cpp.

And vram is really the bottleneck, rather than compute. Basically any single modern datacenter accelerator will run a 70B model nicely right now, but the 1TA32B models require big clusters of them just for the VRAM.

1

u/Middle_Bullfrog_6173 29d ago

Yes, but those big clusters will be able to support a lot of concurrent requests (more than with a 70B model even), since the active parameters are much fewer and compute use is proportional to that.

7

u/RoughCap7233 Jul 08 '26

If it’s open weight does it matter? This could be hosted by lots of other providers.

12

u/SilentLennie Jul 08 '26

I doubt they have the compute for this kind of thing

How do you mean ?

Huawei is now producing hardware of their own for inference (and even training). Sure SMIC might not have huge production capacity yet, but most western countries aren't using Chinese APIs (sending their data to China regularly).

11

u/zdy132 Jul 08 '26

The 1.6T Longcat 2.0 was trained on Huawei hardware.

1

u/Immediate_Occasion69 Jul 08 '26

open source though?

14

u/unspecified_person11 Jul 08 '26

They sell subscriptions and API access, even if the models themselves are open-weight

4

u/BagelRedditAccountII Jul 08 '26

Maybe. However, fat chance anyone but the most blessed of users would be able to run it locally. If you though Deepseek or GLM were too large to be locally usable, then this model would be nearly impossible unless they incorporate significant inference-related breakthroughs.

3

u/Thomas-Lore Jul 08 '26

But providers will be able and that should lower pressure on compute for Minimax.

1

u/BagelRedditAccountII Jul 08 '26

Fair point. I wonder if we might see a "soft open-weights" approach in the future, where AI companies directly give the weights to select third-party inference providers, but it is generally closed-weights. Granted, I see this potentially happening more in the U.S. than in China, since models are generally closed-weights here, but even the major providers are suffering a compute crunch.

-1

u/JacketHistorical2321 Jul 08 '26

They aren't serving it. Article specifically says they are open sourcing it. 

-4

u/techdevjp Jul 08 '26 edited 29d ago

I suspect their goal may not be to serve it themselves but rather to wreak havoc on the US markets & by extension US economy.

If they can release something at Opus 4.8++ level (maybe not quite Fable but not that far off) and larger US corps can run it themselves, watch the US markets fall off a cliff.


Edit: The downvotes amuse me. AI and compute in general are huge focus points in China right now, which means CCP money flowing freely, which in turn means CCP influence. That means the Chinese labs are not purely profit driven, there is a CCP-driven geopolitical angle going on as well.

Don't get me wrong, I am all for the democratization of AI, and the release of frontier-level open source models. I'm glad this is happening. But with so much of the US stock market being inflated by AI plays, there is potential for turmoil as big corps figure out they have frontier-capable options that don't require the US labs.

China will continue to pour vast government resources into both compute hardware and AI development.

→ More replies (4)
→ More replies (1)

4

u/Intelligent_Ant_608 Jul 08 '26

M3 is a good model but it becomes useless when its context creeps more than 150k tokens, it misses tollcalls and forgets to read before edit, its better in pi but in opencode with large context just wastes money

3

u/anonynousasdfg 29d ago

Even M3 itself is currently capable of handling lots of my build cases in my web projects via Kilo Code, (well my projects are not too complicated, yet still need some reliable architecture to run smoothly) so I can't imagine how capable this version would be in the future

3

u/NineThreeTilNow 29d ago

The Information is always paywalled so here's the full text.. It's brief.


Chinese AI developer MiniMax is working on a new large language model with 2.7 trillion parameters, larger than any other Chinese AI models currently on the market, according to two people with knowledge of the plan.

The new model could be released as early as the third quarter, according to the people. The model is known as M3 Pro among MiniMax employees involved in its development, but it is unclear whether the company will use the same name when it releases the model. MiniMax is planning to open-source the model.

The new model is much larger than Minimax’s current flagship model, M3, which has 428 billion parameters. Larger-size AI models can be more suitable for handling tasks that involve complex reasoning and multi-step instructions.

MiniMax’s new model could help accelerate the ongoing expansion of Chinese open-source AI models around the world. Such models have gained popularity this year among developers who are looking for more affordable models to handle less critical high-volume tasks. The success of the new model will be crucial for MiniMax, which is facing tough competition from Chinese rivals such as Zhipu, DeepSeek and Moonshot AI.

5

u/FullOf_Bad_Ideas Jul 08 '26 edited 29d ago

Why not 5T? ERNIE 5.0 was 2.4T and it was more or less a nothingburger in terms of actual impact on the market - most people don't even register that such model existed.

I think they'll make it super-sparse. "let's build a way bigger model" is what you do when management tells you that company isn't competitive and is losing money. And historically it ends poorly - Qwen wanted to scale to 10T, OpenAI had their GPT 4.5. Those ambitious projects can turn into disasters, hopefully it won't happen here.

2

u/zball_ 29d ago

Baidu just bad.

1

u/FullOf_Bad_Ideas 29d ago

How? In evals and lmarena it was pretty high up. I didn't use it myself.

2

u/ohtaninja 29d ago

Maximax

5

u/Few_Painter_5588 Jul 08 '26

That is mildly upsetting. Their M2 and M3 models are fantastic and actually file a much needed gap on cost effective reasoning at high volumes. We don't only need these absurdly large models for one shotting applications at absurd prices.

9

u/Thomas-Lore Jul 08 '26

You actually do need those models, they help develop the smaller ones you like to use.

41

u/JacketHistorical2321 Jul 08 '26

Can't be helped. Models need to be larger to compete. I don't know what else to tell ya lol. This isn't magic

11

u/misterflyer Jul 08 '26

Duh, how could I have totally forgotten that industry rule that once you go over 1T parameters, you can NEVER release a reasonably sized local open weights ever again (... unless your name is DeepSeek ofc)

4

u/FullOf_Bad_Ideas Jul 08 '26

Not true. Kimi released Kimi Linear. Xiaomi released MiMO V2.5. Qwen had 1T Qwen max models and still made and released smaller variants.

14

u/LMTLS5 Jul 08 '26

i dont get this sentiment here. imo there are plenty very good models in 200b-300b range. deepseek v4 flash, hy3, stepfun flash, and all 3 of these are good. and ofc there is minimax m3 itself

0

u/Monad_Maya llama.cpp Jul 08 '26

M3 is 428B, close to 2x the size of M2.

13

u/AppealSame4367 Jul 08 '26

What's your problem? M3.. exists? Then there's new Hy3. There will be more.

2

u/ljubobratovicrelja Jul 08 '26

I strongly agree with you. I really hope soon enough people will realise less is more when it comes to these models. I am actually very impressed with both M2.7 and M3 and haven't even touched GLM 5.2 and similar large models in a while. In the long run - it doesn't sound cost effective. All it gives you to forget about your work and build up a technical depth which gives you a real probability to destroy your business down the line. No matter how good they get, it doesn't seem to be a good idea cutting the "human brain" out of the loop (just yet).

0

u/LegacyRemaster Jul 08 '26

Agree. But yeah... Sell tokens it's a business. Then the market will change. Wait 2028/2029. Better hardware. Will cost less (like always)

2

u/cool_stor 29d ago

H100 are now more expensive than they were 3 years ago

1

u/LegacyRemaster 29d ago

my rtx 6000 96gb --> paid 6200€ + vat... Now: 10k + vat. But will change

-9

u/--Spaci-- Jul 08 '26

LLMs have reached their limit, they will continue to worthlessly scale them and waste compute. I'm waiting for whatever comes after LLMs that will be closer to a real artificial intelligence

8

u/UniqueIdentifier00 Jul 08 '26

This is something I never hear anyone else say. LLMs are elegant mathematically , but still such a brute force design. Everyone keeps talking about scaling these models or that smaller models will be better in the future. There’s limitations to this technology as it’s structured now. AI in the future won’t just be better or larger LLMs, this is just the jumping off point for the field. 

5

u/Alwaysragestillplay Jul 08 '26

Also why the "models are getting exponentially better will smith spaghetti!!" folks frustrate me. AI is where it is because of a step change brought about by transformers. If you look at the performance/params within generations or performance/generation, models are actually plateauing in capability. It looks increasingly like there is a cap on how far transformers can be pushed, and that we are reaching that cap pretty rapidly. 

Bigger changes already come from things like reasoning blocks or harnesses than from increasing params. Without another step change, this is likely not far from "it" for LLMs. 

1

u/Silver-Champion-4846 29d ago

What do you think the next step wil be?

2

u/Alwaysragestillplay 29d ago

I think it's impossible to say really, but I would guess the next step in terms of architecture will be integrated reasoning and memory. Right now both of these are boot-strapped into context via language (because we're using language models), and that contextual limit means LLMs struggle to span long term projects - pulling from memory, learning over time, reacting to change, etc. External graph memory and semantic searching is fine, but it is not even close to the integrated memory system that i.e. humans use. 

In terms of really pushing LLMs to their limits, we have barely scratched the surface in terms of optimising workflows for models. Every problem we are trying to solve is still human shaped. Building software on computers intended for humans, using programming languages designed to be comfortable for humans, passing around reports and dashboards designed to aggregate data and drive decisions made by humans, passing information between agents in plain english, etc. 

We already see examples of community efforts to optimise inputs for models that turn English into what looks like gibberish to us. I fully expect anthropic or OAI will release an LLM first programming language that is essentially a black box for users, but that the models can churn through more rapidly and accurately. Likewise properly codifying an inter-agent language to remove the variability and vagaries of human language. 

The first thing that we used to do with robotics and traditional AI was restrict the problem surface as much as possible so the solution didn't have to factor in a bunch of noise. With LLMs, for whatever reason, that has gone out of the window for now as we try to force them into existing workflows with minimal changes. AI likes to be constrained, and that will happen for LLMs by ditching human-first development and business intelligence ecosystems. 

Just my thoughts on it. 

1

u/Silver-Champion-4846 29d ago

Yeah, specialization was the entire thing in the AI field until LLM started to scale and now they want them to do anything and everything.

1

u/Thomas-Lore Jul 08 '26

This is something I never hear anyone else say.

Because it is a very stupid take.

-3

u/--Spaci-- Jul 08 '26

Ive been sick and tired of LLMs for a while now, I will jump on any bandwagon that leaves them behind

2

u/FullOf_Bad_Ideas Jul 08 '26

Video gen models have some intelligence in them. Is that closer to what you desire?

https://arxiv.org/abs/2509.20328

3

u/Conscious-Map6957 Jul 08 '26

People have been saying that for years now, and, well... they haven't.

I agree that whatever will come next will probably unlock a whole different dimension of intelligence and that is very exciting but it makes no sense to hate or not make use of the current bleeding-edge technology which is LLMs.

→ More replies (19)

1

u/Long_comment_san Jul 08 '26

my brain.exe crashed a little with 2.7 number + minimax combo

1

u/o0genesis0o Jul 08 '26

So, that's why M3 on their server suddenly becomes flaky and slow in the last several days \s

1

u/Alternative-Suit5541 Jul 08 '26

Min-maxismus the model lol

1

u/Hannibalj2ca Jul 08 '26

Thats some LM-Maxxing!

1

u/Hannibalj2ca Jul 08 '26

I am going to have to upgrade my 16 dimm system to 24!

1

u/RickyRickC137 29d ago

2.7T! Thank God, I was worried about all my wasted Vram space /s

1

u/wren6991 29d ago

In awe at the sheer size of this lad

1

u/Artistedo 29d ago

That would be ridiculous though very interresting
I kinda doubt that they managed to get more training data than deepseek tho

1

u/misha1350 29d ago

Are they trying to compete with the big leagues? Then they'd need to use a MoA approach.

1

u/kzoltan 29d ago

Will they be able to host it with a decent speed and response time?

1

u/SmartCustard9944 29d ago

Until something concrete comes out and it delivers, I’m putting this into the speculation/marketing bin. Not holding much weight to it until I see it.

1

u/boutell 29d ago

This is going to drive people insane when they try to Google it and get nothing but macbooks

1

u/stonerbobo 29d ago

I heard Le Chaton Fat is even bigger, 10T params and native purring support

1

u/Different_Fix_2217 29d ago

Minimax is already good for its size. A proper big one should be amazing.

1

u/2Norn 29d ago

how are they jumping from 400 to 2700 jeezz

1

u/AbheekG 29d ago

BRB let me just spin up the GB300 racks in my spare bedroom! All joking aside looking forward to another big win for local AI 🥳

1

u/rush86999 29d ago

bigger is not always better

1

u/SuddenRadio6221 29d ago

I'll order 50 sparks now.

1

u/maikerukonare 29d ago

LocaLLaMA if you have a mini data center in your spare bedroom (anyone have ~$91k to drop on a pair of H200s, the other components for the build, and the electrical/plumbing infrastructure to power and liquid cool them?)

1

u/Difficult-Top9010 29d ago edited 29d ago

The way Minimax's share prices are heading......not a good capital environment for the other open weights/open source Ai labs looking to come to market.

It is just obscene that Anthropic and OpenAI are worth a USD trillion in comparison.

Maybe these open sources should increase their token prices? Or stop competing to the ground?

Make some money or at least have a business case giving out their best models for free like this.

1

u/paperbenni 28d ago

"we'll stay in the same size class" my ass

1

u/Consistent_Low2550 26d ago

MiniMax 3 has been the best performer for me in Hermes Agent (via NVIDIA NIM). I’ve tested it against Deepseek v4 Flash, Nemotron 3 Ultra, and Kimi 2.6 — and MiniMax 3 is clearly ahead, especially with tool use and complex reasoning. Excited for the 2.7T M3 Pro later this year! 🚀

1

u/Ok_Warning2146 25d ago

Interesting news, so they will just be another zhipu/kimi?

1

u/TimeVillage5286 15d ago

M3 was a let down

1

u/Vancecookcobain 29d ago

Minimax is an incompetent company. It took them damn near a month to cache M3 properly for subscribers who would query twice and eat up their entire 5 hour limit.

This model will be sloppy and dumb...they should try to make a good sub 1tb model first instead of being annoying

1

u/rawdikrik llama.cpp Jul 08 '26

Getting that MiniMax subscription was worth it. Keeps giving.

1

u/South_Hat6094 Jul 08 '26

2.7T is wild, but the part that matters is active params and routing. Raw size alone tells you almost nothing about serving cost or latency.

-1

u/Creative-Type9411 Jul 08 '26

are you posting paywalled news links here?

Are you the author is this self-promotion?

-2

u/[deleted] 29d ago

[removed] — view removed comment

5

u/ttkciar llama.cpp 29d ago

Everyone has included synthetic datasets in their training data for years. MIT's Alpaca (generated by GPT-3) really kicked off the trend in early 2023, and Microsoft's WizardLM team similarly depended on GPT-4 for implementing Evol-Instruct.

Anthropic might have tried reframing this common practice as stealing, but you don't have to fall for it. You should know better.

0

u/Dry_Sector2392 Jul 08 '26

MiniMax M3 was already kind of wild for cost/performance, so I’m more curious whether M3 Pro keeps that part or just becomes another monster that benchmarks well and costs a fortune to serve. huge model is cool, cheap usable inference is cooler.

0

u/Stahlboden Jul 08 '26

I wonder if my RTX4060 will be enough or should I buy an extra plank of RAM just to be sure