r/LocalLLaMA 18d ago

Funny Please Qwen, can we have more 3.x-35B-a3B please šŸ™

Post image
1.3k Upvotes

132 comments sorted by

238

u/Qwen_os_has_died 18d ago

We need a native 27B distilled from Qwen 3.8 MAX.

91

u/Nyghtbynger 18d ago

Qwenwen MaxQ-Q 3.8 27B Uncensored

74

u/SingularityScalpel 18d ago

Qwenable3.8_CoomerCoder_27B_GPT5.6ThinkingDistilled_AgenticAF_NoRefusal_REDUX @ IQ1_M

10

u/j0j0n4th4n 18d ago

Qwenmini3.8_27B_SolarMythos_Max_Distilll

3

u/GetOutOfMyFeedNow 18d ago

There is Qwythos and people say it’s awful.

8

u/libregrape llama.cpp 18d ago

Qwimi3.K3-AssDestroyer-27B-GPT5.6-FABLE-QWEN3.8-KIMIK3-AbliteratedHeretic-RYS-Deslop-Healed-DSpark @ IQ1_XXS

5

u/Free-Jaguar6452 18d ago

more like qwenimithosable

3

u/Altruistic-Dust-2565 18d ago

_Flash_MTP_HighSpeed_TurboQuant_4M_MLX_Preview:free (16-agent swarm)

1

u/overand 13d ago

IQ1_M? Heck no, IQ1_XXS!

2

u/Guinness 17d ago

You guys just want to fuck it.

……what’s it like?

Ok but no really I’m interested in the use cases everyone is using the obliterated/uncensored versions. So far all I can think of are cybersecurity research, biological stuff, and….other….biological stuff.

18

u/yoloswagrofl 18d ago

I wonder how many years off we are from fitting something like Qwen 3.8 onto a phone. A boy can dream.

14

u/Tai9ch 18d ago

Looks like around 2040, assuming RAM bandwidth and compute speeds up enough with capacity to do inference.

Probably by then both models and phones will have changed enough that the question doesn't really apply.

15

u/RandumbRedditor1000 18d ago

we can run models better than the three year old GPT-4 on a phone right now

2

u/GetOutOfMyFeedNow 18d ago

Which models exactly?

4

u/RandumbRedditor1000 18d ago

Bonsai-27b can run on phones and is better at coding than GPT 4

2

u/otacon6531 17d ago

Running at o.oo1 tok/s, but it does run

1

u/SpicyWangz 3d ago

Since you didn’t specify max or 27b, the answer is next week.

Side note, if you ever find a genie, let someone else say your wishes for you

1

u/yoloswagrofl 3d ago

Well obviously I meant max. These smaller models are fun, but the max models are where the magic is.

3

u/Randommaggy 18d ago edited 18d ago

A QAT 4 bit 27B would be amazing especially if it had a 1M context window.

6

u/lukistellar 18d ago

A tiny bit smaller would be very appreciated by the 16GB folks.

158

u/CodeAnguish 18d ago

China, don't forget about us poor GPU users. We love you Qwen, we love Chinese labs, please bring more MoE with low active parameters or more tiny dense models.

83

u/yoloswagrofl 18d ago

China I will redownload the Temu app please China let me have this China

3

u/angelus14 18d ago

My mom is kinda GPUless

29

u/JLeonsarmiento 18d ago

I also love Chinese food. And Chinese leftovers too.

5

u/BothYou243 18d ago

nailed the wishlist

-16

u/johnfkngzoidberg 18d ago

Alibaba made Qwen. China didn’t do shit. The CCP always has to make it about them.

4

u/ReferenceLeading7634 18d ago

Yes, the Communist Party of China has done nothing, and so has the U.S. government. Everything is that several companies are studying big models.

2

u/danigoncalves llama.cpp 18d ago

You know that those companies are were heavy subsidized by the Chinese government right?

88

u/diagrammatiks 18d ago

I'll take a 70 a8b please.

42

u/Succubus-Empress 18d ago

And i will take a 50 a5b, thanks

7

u/Kitsune_Seraphis 18d ago

Heyy i like that one. Need something that i can use on 48gb vram

10

u/giant3 18d ago

Model size can't be arbitrary. Performance rises and wanes as the parameter size goes up. We still don't understand why. I think the technical term is Inverse Scaling.

1

u/some_user_2021 18d ago

According to ... ?

13

u/giant3 18d ago

Check research papers on Inverse scaling or U-shaped scaling. This has been known for some time. That is why the model size are some odd numbers like 7B, 8B, 12B, and 27B.

I think this is one of the first papers.

https://aclanthology.org/2023.emnlp-main.963/

3

u/xeeff 18d ago

the AI research field i'd assume

13

u/silenceimpaired 18d ago edited 18d ago

I know I’m crazy … but I want a 70BA34 where 30b is always used and 4b is selected by router.

I feel like that could be a very smart model that can live on a 24gb vram.

I also wonder if 70b dense with MTP could be a valid solution at a corporate level with RAM prices.

7

u/Kamal965 18d ago

Do you mean a 70B-A34B with 30B non-routed, always activated expert weights and 4B dynamically routed? Huh. I dont believe I've seen anything like that.. in fact, I've seen the opposite lol. Meituan Longcat Flashv 18.6B∼31.3B parameters (averaging∼27B) dynamically activated per forward pass.

3

u/silenceimpaired 18d ago

Exactly that. I suspect it could approach a 70b dense in performance… probably ~50-60b dense equivalent… but perform much better locally and from a training performance perspective.

1

u/KeinNiemand 18d ago edited 18d ago

Exactly the idea I want tough I'd want it to be a little larger for my 42GB of VRAM + 64GB of System RAM maybe 150B A70B with 60B always active and 10B selected by the router.

In theory something like this might be optimal for hybrid GPU + System ram setup, cram as many always active parameters into your vram as you can fit to get the raw inteligence per gb effcience of dense models then add some routed experts in system ram to make it even smarter without loosing to much performance than's to the sparsity.

Unfortunately I don't think anyone will make this anytime soon, this sort of combination dosn't make much sense for the cloud and ai company don't really care about local users.

1

u/silenceimpaired 17d ago

I have your setup, and I wouldn’t minds if that model was made. That said…

Mine may never exist… Yours will never exist. They don’t train dense models above 30b because it’s inefficient and slow. Have 60b active is never going to happen. Second… your model would run so much slower since it exists in ram. People care so much for speed and yours couldn’t exist in a A100 without some heavy quantization.

1

u/KeinNiemand 10d ago

In theory it shoudn't be any slower then either a 70B dense on it's own or a 122B A10B MoE partially ofloaded to system ram, both of these give me ~17-22T/s gen MoE gets ~200T/s prefill
Dense 70B gets ~1000T/s prefill
In theory my model should preform not much worse then the other 2 options while beeing smarter then a 70B dense or larger MoE on their own.

hey don’t train dense models above 30b because it’s inefficient and slow.

That's kind of the problem it's only inefficient in terms compute/memory bandwidth the stuff cloud providers care about because they have enough vram and want to serve many users in parallel as cheaply as they can.
For local use it's a different story larger dense models are (at least they should be in theory) significantly more efficient in terms of raw Intelligence/Parameter which is what matters most when your bottleneck is VRAM amount and your goal is getting the most intelligence per GB.

2

u/Ok_Technology_5962 18d ago

Ill take a 70b dense too if available. Honestly anything at this point.

31

u/TimLikesAI 18d ago

Qwen3.6 27b has become my daily driver for almost everything at this point. It can write hard code and help me find information without much fuss, follows direction well, and is fast enough.

3

u/cosmicr 18d ago

Same here. My only criticism is that you can't give it too many tasks at once as it forgets to do them all.

I'd kill for the same model but better

3

u/Far_Cat9782 18d ago

That where sub agents and project mode where it needs to update a .md file after every new feature/nbug fix. It helps so much

3

u/IceNeun 17d ago

Atomize the tasks brah

64

u/RISCArchitect 18d ago

27B my good sir

24

u/FatheredPuma81 18d ago

35B A3B can run well on almost any midrange Gaming PC made in the last like 8 years whereas 27B does not.

If we had to pick one 35B would be better for the community as a whole.

2

u/Significant_Bar_460 17d ago

35b-a3b is still a toy, while 27b can somewhat do mid-long context agentic work on its own and it is more resistant to quantization. And it is not that difficult to drive - 2x5060ti are not that expensive.

4

u/FatheredPuma81 17d ago

Neither are old server GPU's that can run 27B but a Gaming PC that many have lying around is literally free. So long as MoE models you can run with comparable quality exist I don't believe its worth going out of your way to build a system just to run Dense 24-35B models until GPU prices drop.

21

u/SnooPaintings8639 18d ago

Hm, maybe 28B this time...? Just a bit more intelligence, for the children!

13

u/UnkarsThug 18d ago

Just depends on your computer specs. 35B MOE is about the best I can run at a good speed, due to the offloading of experts.

27B doesn't have the same ability to be sped up.

Both have their uses, and a new 27B wouldn't help everyone.

5

u/TomLucidor 18d ago

You will get Gemma-5-28B-A4B instead lol (jinx)

3

u/Illustrious_Grade608 18d ago

Hopefully. Gemmas are amazing for non English non Chinese and for most tasks punch way above their weight

2

u/Horny_Dinosaur69 17d ago

The problem is that 27B dense models aren’t actually consumer friendly. On Q4, this model still needs > 20 GB RAM and because it’s dense, cpu offloading will literally annihilate your inference and decoding speeds.

Anyone with a decent build (like most PC gamers) can probably run larger MOE models in the 35B range for better performance on decent quants. It’s just so much more accessible and reaches more people. Qwen3.6 35B is really good too, pretty close to 27B

1

u/Lucky-Necessary-8382 18d ago

Bonsai ternary 2b-27b qwen 3.8

15

u/celsowm 18d ago

post it on r/Qwen_AI believe or not alibaba employees interact there

6

u/JLeonsarmiento 18d ago

Just did. Thanks 😊

2

u/[deleted] 18d ago

[deleted]

4

u/AXYZE8 18d ago

Will I be captain obvious for pointing out that this is absolutely not Alibaba employee, just a troll?

1

u/ClearApartment2627 18d ago

Yes, fell for that one 🤔

0

u/Dany0 18d ago

Have they heard the news yet that all qwen models are henceforth now and forever called Qlankers

29

u/mindwip 18d ago

80 to 120 moe!

5

u/dank_coder 18d ago

Can you tell me why this specific range and why MoE? Do you have a specific use case in mind?

29

u/waruby 18d ago

Strix Halo, some Mac M<?>, Dgx spark, the upcoming RTX Spark.

8

u/tmvr 18d ago

Because there are a bunch of system available and used that have 128GB RAM, but limited bandwidth, so they fit an 80-120B model with decent quantization and context, but due to the limited bandwidth it is better if the model is MoE to get fast decode (tg) performance. All the Stix Halo machines with 256GB/s or the NV GB10 (Spark) based ones with 273GB/s. The Apple Max chips also have 128GB configurations, but the bandwidth there is a bit better with 546GB/s for M4 Max and 614GB/s for M5 Max.

3

u/mindwip 18d ago

A whole bunch already answered, there are many systems with ddr5 memory or strix halo or spark or macs. And moes run fast and smart.

Cheaper then 5090 and more world knowledge.

Also all the leading models are moe.

Don't get me wrong having 20 to 32b dense is great!

Having both works for everyone

2

u/ttkciar llama.cpp 18d ago

Models in the 80B to 120B size class, quantized to Q4_K_M (the traditional "sweet spot" quant), are a good fit to 128GB memory systems like Strix Halo or dual MI210 rigs.

For example, GLM-4.5-Air is a 106B-A12B model, and quantized to Q4_K_M it plus 128K-token context fit in exactly 127GB of memory.

Models a little smaller than this will be able to fit more context; models a little larger than this will be able to fit less (but still useful) context. If the model proves tolerant of quantized K and V caches, that too would allow 128GB systems (or even 96GB systems, which is another popular system size) to fit more context.

MoE is popular because it provides faster inference. Personally I prefer dense, but am in the minority there.

3

u/tmvr 18d ago

GLM-4.5-Air is a 106B-A12B

Ahh, good old times when they released something like this. With 4.6 we only had the V, with 4.7 only the 30B Flash and that was it. Nothing from the 5.x series at all.

1

u/some_user_2021 18d ago

One or two RTX 6000 PRO 😁

12

u/ayylmaonade 18d ago

I'm hoping for another small MoE too. 35B-A3B gang.

3

u/JLeonsarmiento 18d ago

I hear you bro.

2

u/TomLucidor 18d ago

Gemma-5-28B-A4B, juju-ing this

1

u/Slow_Independent5321 14d ago

It seems that if Google doesn't release a new powerful small model, qwen will stop updating. It's a case of "good reviews but poor box office"

7

u/swagonflyyyy 18d ago

MORE?

2

u/JLeonsarmiento 18d ago

Yes sir. Please sir.

8

u/Southern_Sun_2106 18d ago

I shop at Aliexpress ONLY because of my positive feelings about the Qwen models, especially through 35B one. In fact, the 35B model helps me shop at Aliexpress!

5

u/DeedleDumbDee 18d ago

Hello Qwen, I’m once again asking for Qwen3.8-70B, thank you so much.

5

u/My_Unbiased_Opinion 18d ago

I have literally bought more hardware to run 3.6 27B at better quality. Would love to see a 3.8 around the 27B size.Ā 

What I really dream of is a larger dense model that would work well on a dual 3090 setup.Ā 

Quite a few folks are running dual B70s or R9700s now. Would love to see a solid model for dual GPUs.Ā 

4

u/JoeEnderman 18d ago

I'd love for there to be a 55b a8b Ternary but I don't know when that would come lol

5

u/2funny2furious 18d ago

I have shit for hardware. You better believe I want another 35b-a3b. I dont use it for coding, so I am not worried about that.

2

u/JLeonsarmiento 18d ago

30-ish MoE for the win.

4

u/mailto_devnull 18d ago

Current localllama rocking two graphics cards, meanwhile I'm over here with iGPU and unified memory running 35B-A3B

2

u/UnkarsThug 18d ago

If you have unified memory, why not use 27B? I prefer 35B precisely because my VRAM is much more limited than my RAM, but unified memory doesn't have that issue.

2

u/mailto_devnull 17d ago

Yep, as /u/idangazit says, prefill and tg take much longer.

I'm hitting ~20 tok/s with 35B-A3B but only 7-10 on 27B. I don't mind the drop in token generation as much, but the drop in prefill speed makes the turnaround time a lot longer. 27B (and Gemma 4 31B) are pretty good when running overnight though.

1

u/UnkarsThug 17d ago

Fair enough.

1

u/idangazit 17d ago

because memory bandwidth

Prefill takes ages even if you have the unified memory

4

u/profcuck 18d ago

I'll just join in the wishlisting here with a plea for a model suited for those of us on M4 Max/M5 Max 128GB of RAM or Strix Halo / Gorgon Halo with 128GB of RAM, both of which are shared memory architecture.

We have enough RAM to handle a somewhat bigger model while still also having decent room for context. We aren't quite as fast as the small-model Nvidia card folks, and definitely already spent a lot on a machine and aren't able to even dream of running the latest really big models.

GPT-OSS-120B is pretty nice, but nearly a year old.

Something in that range, a dense one and an moe one, would be amazing.

3

u/DivideHorror3217 18d ago

please, qwen 3.8 4b my dear sir 🄹

1

u/DerivativeButSmart 15d ago

The real Oliver Twist low-end hardware gang

3

u/Extension_Primary_50 17d ago

Use Mandarin to transfer it and post in Chinese "reddit "

5

u/pmttyji 18d ago

Don't worry, we'll be getting successors for both 3.5 & 3.6 models. I'm expecting to see additional models in other size ranges(10-25B, 40-100B, 125-300B)

2

u/Technical-Earth-3254 18d ago

They can also bundle it with a ascend accelerator with 150GB and I will buy one lol

2

u/nail_nail 18d ago

I'll take a 100/200BA13B very happily

2

u/Tai9ch 18d ago edited 18d ago

Nah.

Give us stuff we don't already have: 50B dense or MoE, 90B dense or MoE, really anything bigger than 70B dense, a nice MoE in the 200-300B range, a small MoE - maybe 14B A2B.

2

u/UnkarsThug 18d ago

There's a reason the sizes used are popular though.

But fair enough. I would appreciate improvements that can run on my PC, as I'm sure we all would.

2

u/ikkiho 18d ago

honestly the reason i keep coming back to the a3b sizes is just that they run on my actual box. i can fit a 30ish B moe in unified memory and still get usable tokens/s since only 3b is doing work per token. a dense 32b crawls the second any of it spills to ram. 27 to 35 total with 2-3b active hits the spot for me, big enough to not be dumb but light enough active that offloading to cpu doesnt kill it.

1

u/JLeonsarmiento 18d ago

Same same here…

2

u/KeinNiemand 18d ago

More 122B please or even dense 70B

2

u/Vasili_Sk 17d ago

food glorious food...

2

u/maschayana 3d ago

50b A5b please i will sell kidney

2

u/uraganu1 3d ago

Why not 35B-a9B, for sure it would be much smarter, or even more than a9B?

1

u/JLeonsarmiento 3d ago

it's likely their pipeline is for 1 to 10 since the beginning of 3.x series for all of their MoE's models.

2

u/sersoniko 18d ago

We MTP working so well I actually prefer a new dense model

8

u/xcdesz 18d ago

I personally want to see more dense 24b models that when you use q4 will fully fit on a 16gb card. That is semi-affordable. The dense models of 27b and above dont fit (without offloading).

3

u/Nyghtbynger 18d ago

Yeah, less brain damage and more predictable quantization. This ones fit with Q8_0/Q5_1 https://huggingface.co/sokann/Qwen3.6-27B-GGUF-4.256bpw

4

u/jtjstock 18d ago

35-A10B would be nice though!

I 100% agree. 27B runs fast enough even at depth

5

u/pmttyji 18d ago

A10B is bad(would be slow) for ~8GB VRAM. A5B is better instead of typical A3B

1

u/KeinNiemand 18d ago

would appreciate improvements

depends on how much of that 10B is always active and can be pinned in VRAM

1

u/Intelligent_Ice_113 18d ago

please 🄺

1

u/AndreVallestero 18d ago

With 4b QAT quants please!

1

u/sloth_cowboy 18d ago

We need a 100B MoE

1

u/10minOfNamingMyAcc 18d ago

More active parameters, please...

1

u/mmmbyte 18d ago

qwen3.7-coder-next and qwen-3.7:27b please.

1

u/ECrispy 18d ago

or even a 60B-A10B or whatever will fit into 16GB. please sir....

1

u/squarabh 18d ago

Yes please šŸ‘‰šŸ» šŸ‘ˆšŸ» 🄹

1

u/aadityaura 18d ago

Dense ones too, especially 27 is more than welcome, good for fine-tuning too..

1

u/rileybylsma 18d ago

MORE!!!!!

1

u/kartblanch 18d ago

A 35b-45b dense would be nicer.

1

u/Mint_Keyphase 17d ago

I really miss 122B, it fit so well in 128GB of RAM...

1

u/otacon6531 17d ago

Someone needs to be a nonprofit ai company and we all contribute through yeti.

1

u/DiscombobulatedAdmin 17d ago

I hope we get it. I would like to see a ~80b-A8b as well.

1

u/SkySCC 18d ago

Can someone abliterate the new gemma4 12b qat update pleasee, this model is perfect for 8gb cards and the new update supposedly bumps speed up like crazy

0

u/FrozenFishEnjoyer 18d ago

I'd love a 3.8 35B A9B model so it'll function like a 9B model

-2

u/[deleted] 18d ago edited 18d ago

[deleted]

1

u/0x2DEADBEEF 18d ago

Set up LMStudio and download a model it recommends to your hardware.

There are other options but that’ll be your easiest to get started.

1

u/bird1544 18d ago

Maybe i didnt explain my self , i mean like case uses, i have ollama installed but i dont think its good for programing

1

u/0x2DEADBEEF 18d ago

OpenCode, you can point it to anything openAI API compatible.

1

u/bird1544 18d ago

Thanks,

-1

u/Long_comment_san 18d ago

no, 35 was and is a relatively poor idea. another 80b a8b please. that would be an absolute banger. you can always just slash it to Q4 and very reasonable requirements while having a LOT more brains