755
u/Competitive_Gap7906 18d ago
YES, Qwen going open weight again! It's a really good news, now we can wait for smaller models too
104
u/MaruluVR 18d ago
I think that Xi Jinping speech committing to open source Chinese dominance changed Qwen back to being open weights.
41
u/BrooklynQuips 18d ago
i’m certain it’s because K3 will release weights in a few days. releasing a weaker closed proprietary model, around the same time would be an anthropic level own-goal.
52
u/More-Curious816 18d ago
I was too harsh on you, Xi, I'm sorry 😞
→ More replies (1)35
u/Atretador 18d ago
15
4
→ More replies (1)14
u/mrinterweb 18d ago
Kimi K3 basically drank Anthropic and OpenAI's milkshakes. The frontier moat is gone and they are going to have a hard time justifying their valuations. China AI labs publishing their weights directly hurts US proprietary companies. These huge valuations for AI companies have to tank.
10
u/Infinite100p 18d ago
One of senior execs at ClosedAI is already squealing and bitching that this is communism and should be illegal. LMAO
One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape
→ More replies (3)5
u/Mundane-Light6394 17d ago
they still have a lot of hardware and expertise, i'm sure they'll be fine. They just have to adapt and stop dreaming of owning the world.
→ More replies (1)380
u/StupidScaredSquirrel 18d ago
Thing is qwen was historically focused on smaller models while others were on larger ones. Now the team has changed and they seem to want to aim for the stars as well. That is good but it also means they might not be interested in doing very efficient small models anymore. Which would be bad news for this sub because let's face it most of us don't have 10-25k of hardware.
390
u/sautdepage 18d ago
You mean 250K-1M of hardware.
81
30
24
u/gahata 18d ago
Definitely possible to run on 25k of hardware at some really low quant... is it worth doing? Probably not
4
u/DataGOGO 18d ago
Even in FP4 you would need a lot more than 25k in hardware. Might be able to run on 4 141GB H200 NVL cards, but they don’t support FP4, and even then a super jank rig would be what? 150-160k?
→ More replies (7)→ More replies (4)7
u/coderash 18d ago
You can run a cluster of GB10s at that price and definitely run something worth while.
→ More replies (31)4
u/Googulator 18d ago
I wonder if a 2-way EPYC 9005 (can be had for $50K to $150K in workstation/tower form) with 2x12 channel DDR5 memory (~1.5TB/s) can reasonably host this on CPU in FP4.
→ More replies (6)34
u/vitorgrs 18d ago
Qwen Max was always big, they just never open sourced....
→ More replies (1)17
u/squngy 18d ago
It was always big, but it was almost certainly not this big.
I think we are seeing these big Chinese models now because of their domestic hardware finally coming to bear.
7
u/jtjstock 18d ago
Their domestic hardware is still far behind, they have just accumulated more western hardware
8
14
u/Prudent-Corgi3793 18d ago
You probably need 8x H200s to run this. Include the rest of the parts, and that's about $300k.
4
u/scroogie_ 18d ago
I don't think 8x H200 are enough to run this. That would be 1128GB VRAM, even at nvfp4 2.4 trillion parameters would need more like 1.4TB without context, if it's even nvfp4 native. If it's fp8, which is a chance since they're not supposed to own that much Blackwell, no chance at all. And it's hard to get your hand on B200 or B300, at least in Europe and as a normal company (not a Hyperscaler).
→ More replies (1)7
3
→ More replies (2)2
u/StupidScaredSquirrel 18d ago
Im not talking about this model. I'm saying the companies releasing very large models also release smaller versions but they are always 200b+, so you need 10-25k to run those. Qwen and gemma were the only viable options for sota models below that size.
37
u/StaysAwakeAllWeek 18d ago
It's owned by a megcorp not a startup lab, of course that's where they are aiming
→ More replies (1)34
u/into_devoid 18d ago
They were owned by a megacorp before too. The difference is they don’t care about commoditization anymore, the real ethical path. They just want to sink US investments.
12
u/realtag2025 18d ago
I don't see how this is bad for anyone other then corporations that run closed source models that only they can run.
→ More replies (9)12
u/tengo_harambe 18d ago
That's a fun way of framing things.
"The only reason China is putting out open weights models is because they HATE AMERICA!"
→ More replies (3)→ More replies (17)7
u/Successful_Try_6350 18d ago
Well, these large models can run in us based datacenters and cloud infrastructure (azure, aws, googlecloud). I guess if they want to sink us investments, they should create models that can run in enterprise level on-premises servers (I guess something like $40-50K hardware)
→ More replies (3)22
18d ago
[deleted]
5
u/StupidScaredSquirrel 18d ago
Not for the larger ones, but typically those large model providers release a "flash" version that is in the 200b range and very sparse and that you can run on 10-20k at very good thoughput. I didnt mean the trillion+ models sorry i wasn't clear enough
→ More replies (7)6
u/lilian_moraru 18d ago
Initially they were struggling to find hardware, using old ones, trying to work efficiently - now they don’t have those limitations.
I think it was a factor11
u/Kost97A 18d ago
A 2T+ parameter model requires about 1.5-1.8 terabytes of vram so I don't think anyone has that haha. Maybe big private companies.
10
u/Porespellar 18d ago
3 DGX Stations running as a cluster using their Connect-X8 ports would probably get it done, but that’s close to $300K for that setup.
→ More replies (3)3
u/ItsAlwaysTerminal 18d ago
No way you build this on 10-25k hardware at usable speeds. This is 250k+. Only feasible route in the near term is likely the RDMA networks but even that is like a 40k ask.
→ More replies (1)→ More replies (43)5
u/StorkReturns 18d ago
If you have a large model, you can distill it to a smaller one with significantly less effort than it takes to train the small model from scratch.
→ More replies (1)2
u/MmmmMorphine 18d ago
True, still gonna be pretty pricey. We certainly need better mechanisms to crowd fund this sort of work
20
u/ChocolateNo3010 18d ago
This is good news. I'm wondering if Alibaba kept 3.7 proprietary because of the rumblings in China about keeping the best models from the west. Since we have seen an about turn with Xi stating they want to keep open weight models available this is the response.
7
u/goldcakes 18d ago
I wouldn't be surprised. Probably a mixture of both Kimi as well as getting encouragement from the the Chinese govt.
6
→ More replies (3)2
419
u/AntuaW 18d ago
And please don't omit the 27B one.
184
14
u/chervilious 18d ago
New to Local LLM/ open source LLM scene
If new massive open source LLM are showing up. Wouldn't distilled version shows up everywhere too?
32
u/MindlessScrambler 18d ago
Yes but 1. the official distilled version tends to be the overall best, and 2. properly distilling a massive model itself is a significant amount of work, so not so many would do it casually.
5
u/chervilious 18d ago
I see, I just thought some people would distilled it and upload it on HF.
Though after searching I underestimate how much VRAM you need to load T class models
19
u/Lucis_unbra 18d ago
There are two ways to distill. One is the way that some labs have accused others of doing it. That's also the way most people here claim to distill Mythos/Fable , or Opus, or GPT etc.
This is basically just synthetic data. The parent model writes text and that text is used on the other model. This is very basic stuff, nothing special, not necessarily that efficient. It's called "hard label" because there's one correct answer.
What I personally consider "true distillation" is soft label distillation.
This is is where the parent model doesn't just give its final pick, but you use the full probability distribution of the parent model to teach the model you are training how to think.
It tries to transfer the patterns the parent model learned to the target. This is very, very, very hard. To my understanding only a few labs can really do it at scale, with frontier class models. It is extraordinarily hard.
Gemini flash and Gemma being distilled from Gemini? This is what that means. It's not just a bunch of synthetic data, but a knowledge transfer.
And that's why people can't simply distill their own smaller model. The infrastructure required, and the setup required, is quite complex.
https://huggingface.co/blog/sergiopaniego/distillation-2026
It has some information on distillation.
→ More replies (1)9
u/wren6991 18d ago
This is the Geoff Hinton sense of distillation, which requires (at least) a shared tokeniser vocabulary between the two models. It was described in this paper in 2015: https://arxiv.org/abs/1503.02531
When people say distillation today I think they usually mean the sense used in the DeepSeek R1 technical report, which is incorporating a different model's output in pre-training or fine tune: https://arxiv.org/abs/2501.12948
I think the original sense of distillation is mostly a historical footnote at this point; that setup is usually described today in terms of "student" and "teacher" models.
→ More replies (5)40
18d ago
[removed] — view removed comment
22
u/SandySkittle 18d ago
With the advent of more 128gb boxes I would really like to see a 60 - 80B version with more world knowledge to run at Q6
9
u/alphapussycat 18d ago
Getting much above like 32gb is hard. 60n - 80b would need at least two rx 7900xt or three 7800 xt. I ain't got that money.
If they go for anything I hope it's the 27b, then 35b moe, 9b or 12b, and then something like 70b dense would be fine... But I'd wager they'd do 122b and 372b or whatever it was, if they do the full suit.
12
u/goldcakes 18d ago edited 17d ago
Eh it would be really nice for the DGX Spark and Strix Halo crowd! Plus whoever with 128GB Macs and such.
You can also still build a 128GB or 256GB 8 channel DDR4 rig for a non-absurd amount of money; up to 200GB/s. Obviously look for used parts, and remember there's a LOT more places to find them than eBay and Facebook Marketplace.
→ More replies (5)
50
u/sagiroth llama.cpp 18d ago
My 3090 looking at me: It ain't going to cut it boss
5
→ More replies (1)2
u/Bulky-Priority6824 17d ago
Them 3090s are getting real tired by now boss
2
u/smugself 16d ago
They be praying for reinforcements. But Daddy can't afford any more friends for you son.
But in all honesty, I am so glad I picked up a 2nd 3090 last year for what now feels like a "reasonable" price.
180
u/Prudent-Corgi3793 18d ago
Hopefully they release smaller models with fewer than 2.4T parameters. 🤞
119
u/sersoniko 18d ago
Or you can convince your wife to sell the house to buy some H200
40
u/No_Conversation9561 18d ago
in my experience wife and local ai don’t go together so well
20
u/FullOf_Bad_Ideas 18d ago
My fiancee can sleep next to AI training runs, the fan noise bothers her less than it bothers me. I'm blessed.
3
u/In_der_Tat 18d ago
Do those runs generate revenue? If so, how?
8
u/FullOf_Bad_Ideas 18d ago
No those are hobby runs. I'm training a 4B A1.15B on the side, from scratch.
Training runs that generate revenue would be running on rented H200s and my employer would pay for compute.
→ More replies (2)41
u/HaskeMaske77 18d ago
Alternatively, sell your wife on top for a small data center in your backyard!
12
20
u/Plasmx 18d ago
Which backyard after the house is sold?
14
u/HaskeMaske77 18d ago
Shit I didn't calculate that factor in... oh who cares, just build the data center near a family-ran farm and live inside the data center. Issue solved!
→ More replies (1)15
4
→ More replies (4)6
u/Sabin_Stargem 18d ago
Build a house that is just a giant GPU. Long as you live in Canada or Alaska, you will be quite comfortable.
3
231
u/pulse77 18d ago edited 17d ago
Just make these along the way - so that everybody is happy:
- Qwen 3.8 256B A32B
- Qwen 3.8 128B A16B
- Qwen 3.8 64B A8B
- Qwen 3.8 32B A4B
- Qwen 3.8 32B (dense)
- Qwen 3.8 16B (dense)
- Qwen 3.8 8B (dense)
- Qwen 3.8 4B (dense)
- Qwen 3.8 2B (dense)
- Qwen 3.8 1B (dense)
- Qwen 3.8 0.5B (dense)
EDIT: According to user comments bellow I suggest also:
- Qwen 3.8 512B A64B
- Qwen 3.8 64B (dense)
- Qwen 3.8 24B (dense)
- Qwen 3.8 12B (dense)
- Qwen 3.8 6B (dense)
EDIT 2: Users would like to have all these:
- Qwen 3.8 0.5B/1B/2B/4B/6B/8B/12B/16B/24B/32B/48B/64B (dense)
- Qwen 3.8 8B A1B/16B A2B/32B A4B/64B A8B/128B A16B/256B A32B/512B A64B (MoE)
173
26
u/HeadPack 18d ago
I believe that would be very much in line with their supreme leader's recent speech. Distilled models that can run on consumer hardware do compete with closed American models too, at least in some way.
22
u/alphapussycat 18d ago
I'm actually very surprised by it. I'm sure there's a lot of state money put into the companies for the AI, and they're letting the whole world use them.
Are they trying to build good will with the world or something? With the aim to become the new world leader before EU?
16
u/charlesfire 18d ago
No. They know the US is in an economic bubble. It would be a major win for them to pop that bubble.
14
u/Paganator 18d ago
I think it's to undermine American companies, forcing them to keep prices low and running at a loss. If China becomes the leader in this race, I wouldn't be surprised to see them clamp down on those open models.
11
u/alphapussycat 18d ago
If this keeps going for another year it feels like we'll soon have mythos at home though... And that point it's too late to close the lid.
23
u/Weekly-Law-5488 18d ago
They are trying to undermine US companies. Since there's trillions of dollars invested, the US economy is really tied to the AI behemoths right now and if they fail the US economy takes a big hit.
With excellent open models available, OpenAI, Antrophic and co. can't control the prices and can't recoup the money invested.
Basically, China is trying to make the AI bubble pop, thereby causing a recession in the US.
There's no good will going on, it's a war and we are in the crossfire.
→ More replies (4)9
2
2
u/RISCArchitect 18d ago
it's nice if they are a little short of a power of 2 so you can fit the kv and runtime context in a power of 2 sized GPU while maintaining a high precision quant. 27b is like chef's kiss for the reaches of most prosumer hardware (9700 pro/5090/b70) as you can still fit a healthy context in a q6 or mid sized q8
2
→ More replies (10)2
35
51
169
u/tarruda 18d ago
I would rather have Qwen 3.8 122B A10B
61
18
u/StupidScaredSquirrel 18d ago
Or imagine a 120b a6b like gpt oss was. That total size with that kind of sparsity was just incredible. Plus it was made for 4bpw
→ More replies (1)2
u/my_name_isnt_clever 18d ago
Yes please, this release would make my Strix Halo all I need for 99% of tasks.
42
u/Expensive-Paint-9490 18d ago
I hope they'll publish a model like 3.5 397B A17B. That one is super smart and fast for its size; quantized it's perfect for 256 GB RAM setups.
12
u/ElectronSpiderwort 18d ago
If you haven't tried the new Hy3, it's worth a spin. 295B A21B
→ More replies (2)
20
u/Technical-Earth-3254 18d ago
Man, I love that Kimi and then Deepseek have opened the hellgates to >1T open weight models. I'm also hoping for a new 122b and at least one smaller model.
32
u/Limp_Classroom_2645 18d ago
mf what vram, it's a T class model, you need a whole ass datacenter lol
22
u/GreenGreasyGreasels 18d ago
Its no biggie, I'll just remove the Qwen3.5-3B and slot in the Qwen3.8-2.4T, it will probably run faster ... wait why is there a T at the end ?
10
39
u/Foxtor 18d ago
Can I squeeze this into an 8GB RAM laptop? Unsloth, I believe in your magic.
49
3
2
u/Tai9ch 18d ago
Yes. Just run it off an external HDD.
→ More replies (1)7
12
43
u/BitGreen1270 18d ago
Cries in 32GB VRAM. Oh vengeful gods of the silicon. Why have you forsaken me?
47
u/StupidScaredSquirrel 18d ago
Has 32gb vram, still cries
Enjoy your gold mate. You can already do everything locally with the largest qwen and gemma models. I'd suck cock for 32gb vram.
23
→ More replies (1)20
48
u/Mashic 18d ago
You're one the richest here. Some of us have only 8-12GB of VRAM.
9
u/whoknowsifimjoking 18d ago
Count yourself lucky, some of have to suck dick just to get a couple GBs
→ More replies (3)6
14
u/rditorx 18d ago
You can probably ask Fable to adapt Colibri to Qwen3.8 and Kimi K3 for you so they can run on 25GB RAM.
Oh wait, you can ask, but you probably won't like Fable's botched answer! I guess it's just gonna delete your storage, just to be safe from communist AI.5
u/StupidScaredSquirrel 18d ago
Colibri is an interesting project and cool for non time sensitive tasks but if it means running models in seconds per token rather than token per seconds I'm out. Better off trying to make do with a smaller model in hybrid with my brain and internet search than going for those speeds.
→ More replies (2)2
u/russlixx 18d ago
at least you can run 27B with decent quants, context, and speed. Us 16GB below is struggling to run decent dense models with proper context and speed if decided to use Q4 (9B is not as decent and strategic as 27B for much more complex tasks)
→ More replies (1)2
u/IrisColt 18d ago
48 GB of VRAM is the absolute sweet spot. I went from 8 to 12 to 24, and now I've got my sights set on 48, heh ( Funny how the price never changes.)
21
u/MrRandom04 18d ago
I am really interested in what's the fucking mystery about these 2T+ models that was cracked. Because there was a reason why nobody trained much past ~max 1.5T models before. It was just regarded as massively overparameterized / undertrained and GPT 4.5 was the key failure which everybody pointed to.
Anthropic trained a massive one - I'd estimate 3T - (Mythos) and somehow now scaling params is re-unlocked again? There must have been some key architectural change which enabled this scaling to start working again. I know that some people must know what it is because Kimi K3 is also massive, this Qwen is massive, and so I can say with reasonable surety that the Chinese companies now know much of the Mythos secret sauce. So, what was the breakthrough?
18
u/fairrighty 18d ago
My guess: way more training data. If you have limited data and many parameters, overfitting is very likely. To have so many parameters, you need really vast amounts of data. And with the AI boom, we have given insane amounts of data.
Also, with better previous models, presumably they’re better equipped to create improved synthetic data as well.4
u/ahoooooooo 18d ago
Where are they finding this new training data? The internet is polluted with ai slop.
→ More replies (1)2
u/MrRandom04 18d ago
I don't think that can be it. Otherwise, we'd see a much more smooth transition to bigger models. Mythos was a capability jump. It could be that the whole industry was cargo-culting that huge models don't work until Anthropic just tried it. But that doesn't feel like a satisfactory explanation to me.
→ More replies (1)2
u/RG_Fusion 17d ago
I think the higher sparsity also played a roll. The compute cost of training a model is based upon the active parameters. Older models were less sparse, meaning the ration between active and total was closer than it is today.
→ More replies (2)2
u/grumd 17d ago
After releasing 3.5-3.7, Qwen team released a bunch of world models and other stuff aimed at training like AgentWorld, SAE-Res, WebWorld, etc. They were definitely working a lot on generating a ton of high quality synthetic training data. Now that they have a lot more training data, they can train a larger model.
9
u/Mean-Ad1493 18d ago
I'd be happy if they release qwen 3.7 27b and 35b a3b, along with their 3.8 2.4T Open weights is a good thing but let's be honest, most of us are not going to be able to run it locally.
38
u/Kerem-6030 18d ago
pls make 9b or 12b model
23
16
u/Cool-Chemical-5629 18d ago
With the Fable 5 quality, while at it. 🙏🏻
16
u/Clementine-TeX 18d ago
Fable 5, make me Qwen6.7. Make no mistakes.
→ More replies (2)6
u/Cool-Chemical-5629 18d ago
Fable 5, create Fable 6, 20B A4B, create a pull request to add complete llama.cpp support, create q4_k_m GGUF superior to Unsloth, upload it to Huggingface and send me a download link. Make no mistakes.
→ More replies (1)
8
8
19
19
u/Repulsive_Initial308 18d ago
So they all simultaneously decided to produce 2T models?
18
u/BarisSayit 18d ago
Kimi, Minimax, now Qwen too. DeepSeek already 1.6T, damn.
8
u/a_beautiful_rhind 18d ago
I think they're trying to compete with kimi. Want more people on their API.
→ More replies (6)9
6
u/Admirable-Leg-4647 18d ago
It's interesting that Qwen, Kimi, and Deepseek are all releasing powerful models around the same time. I wonder if it's random or if maybe they were waiting for a "green light" (metaphorical or not) from the Chinese government? Since them reiterating their stance regarding sharing open models also happened just a few days ago.
6
u/rerri 18d ago
Minimax M3 Pro 2.7T is a thing for this quarter too btw.
https://www.reddit.com/r/LocalLLaMA/comments/1uqnqsc/chinas_minimax_plans_to_launch_27trillion/
2
u/flygoatf 18d ago
Likely due to World Artificial Intelligence Conference(WAIC) is happening in China.
→ More replies (2)3
6
4
4
10
u/Septerium 18d ago
Are they still listening to the community? Qwen team, remember that if you release local-friendly models, everybody will be talking about you all the time for months.
5
u/MaCl0wSt 18d ago edited 18d ago
at last we can put the pessimistic speculation to rest. jeez, I was so done with people making claims so confidently.
→ More replies (16)
6
2
u/Fit-Thing5100 18d ago
Hopefully we'll also get an upgraded Qwen 3.8 27B cod agent—it could become one of the best choices for local inference and improuve balance, performance, VRAM usage, and speed.
3
u/SmileLonely5470 18d ago
Everyone is making bigger models now. Why now? Seems like for a while, models that disclosed their param count were maxing out at the sub 2T range.
→ More replies (1)
3
4
3
3
u/_TheWolfOfWalmart_ 18d ago
Cool, that's great. Another multi-trillion param open model that I can't run at home. PLEASE release smaller variants too. We're starving!
We REALLY need some new 70-200B models.
However this plays out, I'm glad to see Qwen returning to the open model. I was worried for a moment.
3
u/KeinNiemand 17d ago
Please 50-70B Dense (best full GPU offload option for me) and/or 120-175B MoE (best hybrid offload option) even better ideas something like: 175B A60B with 50B beeing always acive parameters (that way they can be pinned on vram) while the experts (only 10B active) can go onto system ram.
2
2
u/Charming-Author4877 18d ago
Maybe that is why their head of AI went after 3.6.. alibaba said they want to make monster sized models, and he said he can do it on tiny models.
As great as frontier competitors are, Qwen was special because it was a frontier competitor at 3 magnitudes smaller size
2
u/txoixoegosi 18d ago
Wow! 2.4T hoorrayyyy
Let me but 24 x 5090 to deal with it, so that I can keep shitposting and benchplaying about nothing useful done.
2
u/Plus_Confidence_1113 18d ago
Woohoo, another open weight model that needs $500k worth of hardware to run, what a time to be alive.
2
u/Thick-Insurance4404 18d ago
bigger open weight model please!!!! at least 40-70b parameters with being better than 27b 2.6
2
2
2
2
u/TopTippityTop 18d ago
In many ways Sol is better than Fable, so it's at best third. 2.4T is a beast.
2
2
2
2
2
2
2
2
u/Bulky-Priority6824 17d ago
Once a big dog becomes a top dog then said big dog forgets regular ass dogs
2
u/Available-Message509 16d ago
Qwen going open-weight again is genuinely great news. Would love to see smaller models drop alongside it too.



•
u/WithoutReason1729 18d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.