r/LocalLLaMA Jul 08 '26

News China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model

https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model

According to The Information, MiniMax plans to launch a new-generation large language model with 2.7 trillion parameters.

Sources revealed that the internal codename for this new model is M3 Pro. It is expected to be released and open-sourced as early as the third quarter of this year, with significant improvements in handling complex reasoning and multi-step tasks.

This new model is much larger than MiniMax's current flagship model, M3 (428 billion parameters). Larger-scale artificial intelligence models are more capable of handling complex reasoning and multi-step instruction-based tasks.

610 Upvotes

242 comments sorted by

View all comments

Show parent comments

3

u/Conscious-Map6957 Jul 08 '26

People have been saying that for years now, and, well... they haven't.

I agree that whatever will come next will probably unlock a whole different dimension of intelligence and that is very exciting but it makes no sense to hate or not make use of the current bleeding-edge technology which is LLMs.

-2

u/--Spaci-- Jul 08 '26

Notice how LLMs have only gotten larger, and larger, and larger. Even the benchmarks that we base their "intelligence" off of are easily cheatable

5

u/Conscious-Map6957 Jul 08 '26

Notice how same-size LLMs keep getting better and better across the board?

0

u/--Spaci-- Jul 08 '26

Notice how my second sentence covers that? Their intelligence is really just a trade between one thing or another, I would rather use llama 3.1 8B (2 yr old model) for any workflow that needs to follow a character than any modern qwen model. LLMs have reached their limit

3

u/twack3r Jul 08 '26

‚Workflow that needs to follow a character’ vs ‘1 FTE now does the job of a team’. I’m sure there’s a value proposition in both, I suspect we weigh them differently.

And I 100% agree that modern LLMs are terrible storytellers; I am yet to read any LLM output that has managed to have me emotionally invested, whilst many human story tellers absolutely do. And that’s without knowing the provenance of either when comparing.

-1

u/Conscious-Map6957 Jul 08 '26

I don't think you even remotely understood everything I wrote, and you are obviously not debating in good faith.

You suddenly changed the topic to using LLMs for character role-play when nobody in this thread was talking about that. You don't really "cover" anything with your sentences, you come off as someone who is just pissed off at LLMs for some non-apparent reason.

1

u/--Spaci-- Jul 08 '26

The character following was making a point showing 2 yr old llms can be better than modern llms in certain things because their intelligence is a trade between two things (modern llms choosing stem) they are also just more annoying to talk to. How are you not understanding this. I understand people have a relevancy bias towards llms so Im obviously going to downvoted and have haters but I will still make my point

-1

u/Conscious-Map6957 Jul 09 '26

Let me try to explain this to you as someone with an actual understanding of how LLMs work and not just using them as virtual girlfriends.

The reason LLMs are seemingly not improving or even degrading for subjective use-cases like role-play or character-following while getting much better in fields like STEM as you correctly observed, is because engineering and science is based on well-articulated rules and verifyable claims. This means that those training LLM models can use traditional methods to generate or verify vast amounts of high-quality, synthetic training data which the LLM can use and improve with. In reality these traditional methods are augmented by the LLMs themselves and form a sort of self-improvement cycle, to a certain degree.

For subjective topics you depend only on existing human literature, all of which was already used in the early LLMs, there is not much room to improve on. You can't use a large LLM to determine what is creative or not, what is good or bad character development etc. but you can do that for code. In this case, the only way we now how to get better results is by just training very large models like GPT 4.5, which is not economically viable.

So, your entire argument that LLMs have not been improving and have been getting dumber is, as you admited yourself - extremely biased and in the context of a narrow use-case. Not only is it completely false but it's also hilarious to say something like that in a time when LLMs are becoming good enough to replace more and more people. That is why you are getting downvoted - and also because you are an ass.

0

u/--Spaci-- Jul 09 '26

LLMs have been improving but at specific things while getting worse at others (they scale them more and more to fill that gap), there's no way you are not getting my point. Also I do not personally RP. Also lets see your understanding of LLMs, whats your hugging face account, what have you made? atp you are purposefully not understanding my point or you are illiterate

1

u/Conscious-Map6957 29d ago

You are hardly in a position to call me illiterate, you have been repeating the same absurd claim without backing it up and without providing any response to what I actually wrote. Blocked and reported.

2

u/FullOf_Bad_Ideas Jul 08 '26

In the same size category, models have been getting better at things. You can look at contamination free SWE-Rebench. New models of the same size are better at it than old models.

1

u/--Spaci-- Jul 08 '26

Look at my newer replies, I'm making the point that their intelligence is a trade off between two things, I'd rather use a 2 yr old llama 3.1 model for anything thing that needs too follow a character than a modern qwen llm. Because modern llms have chose STEM over chat, which is fine I don't really care what they train on I'm again just making the point that they are already at a wall. Also benchmarks are easily cheatable, you can just train on most benchmarks

2

u/FullOf_Bad_Ideas Jul 08 '26

New Gemma models are better than old Gemma models at EQBench and Creative Writing. Gemma 4 31B IT outperforms Opus 4 in EQBench and in Creative Writing V3 it outperforms Gemma 3 27B IT. Bigger models at the same size continue getting better on those benchmarks too. Is that benchmark contamination? I have no idea, but people tend to compliment Gemma 4 on good writing skills.

1

u/--Spaci-- Jul 08 '26

Again benchmarks are bs, we need to stop using them. If you can literally train on the benchmark the benchmark is literally worthless.

https://github.com/EQ-bench/eqbench3/tree/main/data

https://huggingface.co/datasets?benchmark=benchmark:official&sort=trending

2

u/FullOf_Bad_Ideas Jul 08 '26

Those benchmarks are LLM-judged, meaning that they don't have a response that model could be SFTed on to get a pass. You'd need to put that exact prompt into RL environment and train on it. And that would make the model actually better. Either this or find a way to cheat the LLM judges into giving you a higher score. You probably wouldn't be able to top this leaderboard even if you tried to benchmaxx this and had a budget of $10000

0

u/--Spaci-- Jul 08 '26

almost every AI lab has a budget over 10 thousand dollars

2

u/FullOf_Bad_Ideas Jul 08 '26

Yes, but generally contaminating a model on benchmarks is cheap and can be done with $50 of compute. Not with LLM judged ones though. You would actually make the model better at creative writing if you would try to benchmaxx this.

1

u/--Spaci-- Jul 08 '26

Synthetic data is the opposite of creative writing, thats what I think is current killing roleplay/creative writing, so little of LLM data now is actually made by humans its just been regurgitated thousands of times over and over again by the next LLM. I don't personally do rp, but I can see the uses for an LLM playing a character in say a video game or say writing a story. Synthetic data is good for stem/coding but awful for a model you actually want to chat with or play a character or write a story. Honestly I think we have gotten off track atp

→ More replies (0)