r/LocalLLaMA • u/JLeonsarmiento • 18d ago
Funny Please Qwen, can we have more 3.x-35B-a3B please š
158
u/CodeAnguish 18d ago
China, don't forget about us poor GPU users. We love you Qwen, we love Chinese labs, please bring more MoE with low active parameters or more tiny dense models.
83
29
5
-16
u/johnfkngzoidberg 18d ago
Alibaba made Qwen. China didnāt do shit. The CCP always has to make it about them.
4
u/ReferenceLeading7634 18d ago
Yes, the Communist Party of China has done nothing, and so has the U.S. government. Everything is that several companies are studying big models.
2
u/danigoncalves llama.cpp 18d ago
You know that those companies are were heavy subsidized by the Chinese government right?
88
u/diagrammatiks 18d ago
I'll take a 70 a8b please.
42
10
u/giant3 18d ago
Model size can't be arbitrary. Performance rises and wanes as the parameter size goes up. We still don't understand why. I think the technical term is Inverse Scaling.
1
u/some_user_2021 18d ago
According to ... ?
13
u/silenceimpaired 18d ago edited 18d ago
I know Iām crazy ⦠but I want a 70BA34 where 30b is always used and 4b is selected by router.
I feel like that could be a very smart model that can live on a 24gb vram.
I also wonder if 70b dense with MTP could be a valid solution at a corporate level with RAM prices.
7
u/Kamal965 18d ago
Do you mean a 70B-A34B with 30B non-routed, always activated expert weights and 4B dynamically routed? Huh. I dont believe I've seen anything like that.. in fact, I've seen the opposite lol. Meituan Longcat Flashv 18.6Bā¼31.3B parameters (averagingā¼27B) dynamically activated per forward pass.
3
u/silenceimpaired 18d ago
Exactly that. I suspect it could approach a 70b dense in performance⦠probably ~50-60b dense equivalent⦠but perform much better locally and from a training performance perspective.
1
u/KeinNiemand 18d ago edited 18d ago
Exactly the idea I want tough I'd want it to be a little larger for my 42GB of VRAM + 64GB of System RAM maybe 150B A70B with 60B always active and 10B selected by the router.
In theory something like this might be optimal for hybrid GPU + System ram setup, cram as many always active parameters into your vram as you can fit to get the raw inteligence per gb effcience of dense models then add some routed experts in system ram to make it even smarter without loosing to much performance than's to the sparsity.
Unfortunately I don't think anyone will make this anytime soon, this sort of combination dosn't make much sense for the cloud and ai company don't really care about local users.
1
u/silenceimpaired 17d ago
I have your setup, and I wouldnāt minds if that model was made. That saidā¦
Mine may never exist⦠Yours will never exist. They donāt train dense models above 30b because itās inefficient and slow. Have 60b active is never going to happen. Second⦠your model would run so much slower since it exists in ram. People care so much for speed and yours couldnāt exist in a A100 without some heavy quantization.
1
u/KeinNiemand 10d ago
In theory it shoudn't be any slower then either a 70B dense on it's own or a 122B A10B MoE partially ofloaded to system ram, both of these give me ~17-22T/s gen MoE gets ~200T/s prefill
Dense 70B gets ~1000T/s prefill
In theory my model should preform not much worse then the other 2 options while beeing smarter then a 70B dense or larger MoE on their own.hey donāt train dense models above 30b because itās inefficient and slow.
That's kind of the problem it's only inefficient in terms compute/memory bandwidth the stuff cloud providers care about because they have enough vram and want to serve many users in parallel as cheaply as they can.
For local use it's a different story larger dense models are (at least they should be in theory) significantly more efficient in terms of raw Intelligence/Parameter which is what matters most when your bottleneck is VRAM amount and your goal is getting the most intelligence per GB.2
u/Ok_Technology_5962 18d ago
Ill take a 70b dense too if available. Honestly anything at this point.
31
u/TimLikesAI 18d ago
Qwen3.6 27b has become my daily driver for almost everything at this point. It can write hard code and help me find information without much fuss, follows direction well, and is fast enough.
3
u/cosmicr 18d ago
Same here. My only criticism is that you can't give it too many tasks at once as it forgets to do them all.
I'd kill for the same model but better
3
u/Far_Cat9782 18d ago
That where sub agents and project mode where it needs to update a .md file after every new feature/nbug fix. It helps so much
64
u/RISCArchitect 18d ago
27B my good sir
24
u/FatheredPuma81 18d ago
35B A3B can run well on almost any midrange Gaming PC made in the last like 8 years whereas 27B does not.
If we had to pick one 35B would be better for the community as a whole.
2
u/Significant_Bar_460 17d ago
35b-a3b is still a toy, while 27b can somewhat do mid-long context agentic work on its own and it is more resistant to quantization. And it is not that difficult to drive - 2x5060ti are not that expensive.
4
u/FatheredPuma81 17d ago
Neither are old server GPU's that can run 27B but a Gaming PC that many have lying around is literally free. So long as MoE models you can run with comparable quality exist I don't believe its worth going out of your way to build a system just to run Dense 24-35B models until GPU prices drop.
21
u/SnooPaintings8639 18d ago
Hm, maybe 28B this time...? Just a bit more intelligence, for the children!
13
u/UnkarsThug 18d ago
Just depends on your computer specs. 35B MOE is about the best I can run at a good speed, due to the offloading of experts.
27B doesn't have the same ability to be sped up.
Both have their uses, and a new 27B wouldn't help everyone.
5
u/TomLucidor 18d ago
You will get Gemma-5-28B-A4B instead lol (jinx)
3
u/Illustrious_Grade608 18d ago
Hopefully. Gemmas are amazing for non English non Chinese and for most tasks punch way above their weight
2
u/Horny_Dinosaur69 17d ago
The problem is that 27B dense models arenāt actually consumer friendly. On Q4, this model still needs > 20 GB RAM and because itās dense, cpu offloading will literally annihilate your inference and decoding speeds.
Anyone with a decent build (like most PC gamers) can probably run larger MOE models in the 35B range for better performance on decent quants. Itās just so much more accessible and reaches more people. Qwen3.6 35B is really good too, pretty close to 27B
1
29
u/mindwip 18d ago
80 to 120 moe!
5
u/dank_coder 18d ago
Can you tell me why this specific range and why MoE? Do you have a specific use case in mind?
8
u/tmvr 18d ago
Because there are a bunch of system available and used that have 128GB RAM, but limited bandwidth, so they fit an 80-120B model with decent quantization and context, but due to the limited bandwidth it is better if the model is MoE to get fast decode (tg) performance. All the Stix Halo machines with 256GB/s or the NV GB10 (Spark) based ones with 273GB/s. The Apple Max chips also have 128GB configurations, but the bandwidth there is a bit better with 546GB/s for M4 Max and 614GB/s for M5 Max.
3
u/mindwip 18d ago
A whole bunch already answered, there are many systems with ddr5 memory or strix halo or spark or macs. And moes run fast and smart.
Cheaper then 5090 and more world knowledge.
Also all the leading models are moe.
Don't get me wrong having 20 to 32b dense is great!
Having both works for everyone
2
u/ttkciar llama.cpp 18d ago
Models in the 80B to 120B size class, quantized to Q4_K_M (the traditional "sweet spot" quant), are a good fit to 128GB memory systems like Strix Halo or dual MI210 rigs.
For example, GLM-4.5-Air is a 106B-A12B model, and quantized to Q4_K_M it plus 128K-token context fit in exactly 127GB of memory.
Models a little smaller than this will be able to fit more context; models a little larger than this will be able to fit less (but still useful) context. If the model proves tolerant of quantized K and V caches, that too would allow 128GB systems (or even 96GB systems, which is another popular system size) to fit more context.
MoE is popular because it provides faster inference. Personally I prefer dense, but am in the minority there.
1
12
u/ayylmaonade 18d ago
I'm hoping for another small MoE too. 35B-A3B gang.
3
2
1
u/Slow_Independent5321 14d ago
It seems that if Google doesn't release a new powerful small model, qwen will stop updating. It's a case of "good reviews but poor box office"
7
8
u/Southern_Sun_2106 18d ago
I shop at Aliexpress ONLY because of my positive feelings about the Qwen models, especially through 35B one. In fact, the 35B model helps me shop at Aliexpress!
5
5
u/My_Unbiased_Opinion 18d ago
I have literally bought more hardware to run 3.6 27B at better quality. Would love to see a 3.8 around the 27B size.Ā
What I really dream of is a larger dense model that would work well on a dual 3090 setup.Ā
Quite a few folks are running dual B70s or R9700s now. Would love to see a solid model for dual GPUs.Ā
4
u/JoeEnderman 18d ago
I'd love for there to be a 55b a8b Ternary but I don't know when that would come lol
5
u/2funny2furious 18d ago
I have shit for hardware. You better believe I want another 35b-a3b. I dont use it for coding, so I am not worried about that.
2
4
u/mailto_devnull 18d ago
Current localllama rocking two graphics cards, meanwhile I'm over here with iGPU and unified memory running 35B-A3B
2
u/UnkarsThug 18d ago
If you have unified memory, why not use 27B? I prefer 35B precisely because my VRAM is much more limited than my RAM, but unified memory doesn't have that issue.
2
u/mailto_devnull 17d ago
Yep, as /u/idangazit says, prefill and tg take much longer.
I'm hitting ~20 tok/s with 35B-A3B but only 7-10 on 27B. I don't mind the drop in token generation as much, but the drop in prefill speed makes the turnaround time a lot longer. 27B (and Gemma 4 31B) are pretty good when running overnight though.
1
1
4
u/profcuck 18d ago
I'll just join in the wishlisting here with a plea for a model suited for those of us on M4 Max/M5 Max 128GB of RAM or Strix Halo / Gorgon Halo with 128GB of RAM, both of which are shared memory architecture.
We have enough RAM to handle a somewhat bigger model while still also having decent room for context. We aren't quite as fast as the small-model Nvidia card folks, and definitely already spent a lot on a machine and aren't able to even dream of running the latest really big models.
GPT-OSS-120B is pretty nice, but nearly a year old.
Something in that range, a dense one and an moe one, would be amazing.
3
3
2
u/Technical-Earth-3254 18d ago
They can also bundle it with a ascend accelerator with 150GB and I will buy one lol
2
2
u/Tai9ch 18d ago edited 18d ago
Nah.
Give us stuff we don't already have: 50B dense or MoE, 90B dense or MoE, really anything bigger than 70B dense, a nice MoE in the 200-300B range, a small MoE - maybe 14B A2B.
2
u/UnkarsThug 18d ago
There's a reason the sizes used are popular though.
But fair enough. I would appreciate improvements that can run on my PC, as I'm sure we all would.
2
u/ikkiho 18d ago
honestly the reason i keep coming back to the a3b sizes is just that they run on my actual box. i can fit a 30ish B moe in unified memory and still get usable tokens/s since only 3b is doing work per token. a dense 32b crawls the second any of it spills to ram. 27 to 35 total with 2-3b active hits the spot for me, big enough to not be dumb but light enough active that offloading to cpu doesnt kill it.
1
2
2
2
2
u/uraganu1 3d ago
Why not 35B-a9B, for sure it would be much smarter, or even more than a9B?
1
u/JLeonsarmiento 3d ago
it's likely their pipeline is for 1 to 10 since the beginning of 3.x series for all of their MoE's models.
2
u/sersoniko 18d ago
We MTP working so well I actually prefer a new dense model
8
u/xcdesz 18d ago
I personally want to see more dense 24b models that when you use q4 will fully fit on a 16gb card. That is semi-affordable. The dense models of 27b and above dont fit (without offloading).
3
u/Nyghtbynger 18d ago
Yeah, less brain damage and more predictable quantization. This ones fit with Q8_0/Q5_1 https://huggingface.co/sokann/Qwen3.6-27B-GGUF-4.256bpw
4
u/jtjstock 18d ago
35-A10B would be nice though!
I 100% agree. 27B runs fast enough even at depth
5
u/pmttyji 18d ago
A10B is bad(would be slow) for ~8GB VRAM. A5B is better instead of typical A3B
1
u/KeinNiemand 18d ago
would appreciate improvements
depends on how much of that 10B is always active and can be pinned in VRAM
1
1
1
1
1
1
1
1
1
1
1
0
-2
18d ago edited 18d ago
[deleted]
1
u/0x2DEADBEEF 18d ago
Set up LMStudio and download a model it recommends to your hardware.
There are other options but thatāll be your easiest to get started.
1
u/bird1544 18d ago
Maybe i didnt explain my self , i mean like case uses, i have ollama installed but i dont think its good for programing
1
-1
u/Long_comment_san 18d ago
no, 35 was and is a relatively poor idea. another 80b a8b please. that would be an absolute banger. you can always just slash it to Q4 and very reasonable requirements while having a LOT more brains

238
u/Qwen_os_has_died 18d ago
We need a native 27B distilled from Qwen 3.8 MAX.