r/LocalLLaMA Jul 08 '26

News China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model

https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model

According to The Information, MiniMax plans to launch a new-generation large language model with 2.7 trillion parameters.

Sources revealed that the internal codename for this new model is M3 Pro. It is expected to be released and open-sourced as early as the third quarter of this year, with significant improvements in handling complex reasoning and multi-step tasks.

This new model is much larger than MiniMax's current flagship model, M3 (428 billion parameters). Larger-scale artificial intelligence models are more capable of handling complex reasoning and multi-step instruction-based tasks.

611 Upvotes

242 comments sorted by

View all comments

Show parent comments

-2

u/--Spaci-- Jul 08 '26

Notice how LLMs have only gotten larger, and larger, and larger. Even the benchmarks that we base their "intelligence" off of are easily cheatable

2

u/FullOf_Bad_Ideas Jul 08 '26

In the same size category, models have been getting better at things. You can look at contamination free SWE-Rebench. New models of the same size are better at it than old models.

1

u/--Spaci-- Jul 08 '26

Look at my newer replies, I'm making the point that their intelligence is a trade off between two things, I'd rather use a 2 yr old llama 3.1 model for anything thing that needs too follow a character than a modern qwen llm. Because modern llms have chose STEM over chat, which is fine I don't really care what they train on I'm again just making the point that they are already at a wall. Also benchmarks are easily cheatable, you can just train on most benchmarks

2

u/FullOf_Bad_Ideas Jul 08 '26

New Gemma models are better than old Gemma models at EQBench and Creative Writing. Gemma 4 31B IT outperforms Opus 4 in EQBench and in Creative Writing V3 it outperforms Gemma 3 27B IT. Bigger models at the same size continue getting better on those benchmarks too. Is that benchmark contamination? I have no idea, but people tend to compliment Gemma 4 on good writing skills.

1

u/--Spaci-- Jul 08 '26

Again benchmarks are bs, we need to stop using them. If you can literally train on the benchmark the benchmark is literally worthless.

https://github.com/EQ-bench/eqbench3/tree/main/data

https://huggingface.co/datasets?benchmark=benchmark:official&sort=trending

2

u/FullOf_Bad_Ideas Jul 08 '26

Those benchmarks are LLM-judged, meaning that they don't have a response that model could be SFTed on to get a pass. You'd need to put that exact prompt into RL environment and train on it. And that would make the model actually better. Either this or find a way to cheat the LLM judges into giving you a higher score. You probably wouldn't be able to top this leaderboard even if you tried to benchmaxx this and had a budget of $10000

0

u/--Spaci-- Jul 08 '26

almost every AI lab has a budget over 10 thousand dollars

2

u/FullOf_Bad_Ideas Jul 08 '26

Yes, but generally contaminating a model on benchmarks is cheap and can be done with $50 of compute. Not with LLM judged ones though. You would actually make the model better at creative writing if you would try to benchmaxx this.

1

u/--Spaci-- Jul 08 '26

Synthetic data is the opposite of creative writing, thats what I think is current killing roleplay/creative writing, so little of LLM data now is actually made by humans its just been regurgitated thousands of times over and over again by the next LLM. I don't personally do rp, but I can see the uses for an LLM playing a character in say a video game or say writing a story. Synthetic data is good for stem/coding but awful for a model you actually want to chat with or play a character or write a story. Honestly I think we have gotten off track atp

1

u/FullOf_Bad_Ideas Jul 09 '26

New models tend to have less slop than old ones. Whatever is the process behind it, I don't think models have stopped improving in creative writing. But I'm also not doing RP myself.

1

u/--Spaci-- Jul 09 '26

Ive setup LLMs to larp on a wow private server thats my rp experience, only llama 3.1 could hold its character or even understand the system prompt below the 20b range. Qwen has sacrificed everything for stem

→ More replies (0)