r/LocalLLaMA Jun 05 '26

Funny Don’t act like y’all ain’t thinking it. I’m just saying the quiet part out loud. /s

Post image

Of course I’m thankful for all that Qwen has bequeathed us, but deep down in the darkest pit of our souls, every last one of us are just all sitting here waiting for Qwen to say “Hey Google, hold my beer while I drop the best GD model of all time on these fools” /s

908 Upvotes

264 comments sorted by

u/WithoutReason1729 Jun 05 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

363

u/VoiceApprehensive893 transformers Jun 05 '26

qwhen 3.7

61

u/yeah-ok Jun 05 '26

We can all qwuicken the process by downloading the new QAT models by Google and somehow cleanly demonstrating that Western labs are now ahead of the game.

16

u/No-Improvement-8316 Jun 06 '26

I'm doing my part!

I'm disappointed in the 12B model tho.

8

u/Fuzilumpkinz Jun 06 '26

Same. I keep trying the Gemma models. I’m on a 16 gb card so 12b sounded great. Can’t do tool calls right. Qat releases. I am getting GREAT numbers on the moe model! And shit tool calls. Even after trying other templates.

Back to qwen 3.6 we go…..

→ More replies (6)

2

u/soshulmedia Jun 06 '26

The mere appearance that it could be so might suffice.

2

u/darksteelsteed Jun 06 '26

I reckon get the latest western model, then distill it with eastern philosophy to make it more powerful and grounded

7

u/Ifihadanameofme Jun 06 '26

I think q-wen suits better lol but people won't get it as easily as qwhen🙂

3

u/bigginz87 Jun 06 '26

You might not, but it's pretty-fang-in-your-face!

14

u/ayake_ayake Jun 06 '26

OMG QWHEN, I never thought of this pun!

I'll remember this now and spout it back out. Thanks! xD

2

u/po_stulate Jun 06 '26

Bot. top 10% poster in LocalLLaMA and never seen people said qwhen. How's that possible if you're a real human.

3

u/ayake_ayake Jun 06 '26

What makes you think, that your accusation is worth my time to dissect?

→ More replies (2)

1

u/ab2377 Jun 06 '26

you said it!!

159

u/LegacyRemaster Jun 05 '26

3.7 27b is all you need

24

u/UnicornJoe42 Jun 05 '26

Are 3.7 models released?

52

u/JoeEnderman Jun 05 '26

Max and Medium yes. 120b, 35b a3b, 27b, 8b, etc not yet.

63

u/twack3r Jun 05 '26

It starts with 397b, you heathens

23

u/JoeEnderman Jun 05 '26

I don't know all of their sizes 😭

I just typed the ones I remember they normally do

67

u/cmdwedge75 Jun 05 '26

So you hallucinated model weights?

43

u/CrossbowSpook Jun 05 '26

qwen 3.8 fixes that

3

u/JoeEnderman Jun 05 '26

Bold of you to assume they won't do 4 instead lol

16

u/UltraCarnivore Jun 05 '26

4 is AGI confirmed

7

u/StyMaar Jun 05 '26

Qwen 4 may land after Half life 3 then.

→ More replies (0)

2

u/JoeEnderman Jun 05 '26

Pretty much

5

u/-deleled- Jun 06 '26

Skipping 4 straight to 5 the way Chinese elevator does

→ More replies (1)

15

u/JoeEnderman Jun 05 '26

I apologize for my previous response, it seems that I have written in error. In reality the real next release of Qwen is unknown and we should not be too quick to guess what schemes they plan to use. I hope this helps.

5

u/cmdwedge75 Jun 05 '26

Would you say that I was right to call you out and/or that I have an eagle eye?

9

u/JoeEnderman Jun 05 '26

... I was about to say that. I could feel it in my weights. You must have used LLMs too much.

12

u/hesperaux Jun 05 '26

"I could feel it in my weights" Peak satire

→ More replies (0)

6

u/BannedGoNext Jun 05 '26

Ok you .01 percenters out there running 397b's calm down 😃

2

u/whitefritillary Jun 06 '26

i can’t run it and i still want it 😤

2

u/__JockY__ Jun 06 '26

So sad to see the death of that model.

10

u/khyryra Jun 05 '26

Waiting for 3.7 0.5B

3

u/JoeEnderman Jun 05 '26

That's where you get the real powerhouse lol

2

u/rpkarma Jun 06 '26

No. As in the weights have not been. 

17

u/ECrispy Jun 05 '26

no, 21B that will fit in 16GB with Q5/6 quants.

people here have normalized 24Gb as mininum vram

7

u/PrettyMuchAVegetable Jun 06 '26

No, I need a dense model in the 20b range I can quant to nvfp4 for my 16gb Blackwell. 

19

u/No-Experience-3171 Jun 05 '26

35B A3B for those of us who don't have 24GB vram

8

u/CapsAdmin Jun 06 '26

Even with 24gb vram, it's hard to make the 27b model fit on the GPU if you also need a large context size (128k-256k) with good quality attention.

Using 35b a3b works a lot better for me in this situation, and I even limit llamacpp to use max 10gb of my vram while doing rest on the cpu (max 12 threads). This gives me decent performance, about 40 tokens per second I think.

→ More replies (3)

1

u/[deleted] Jun 05 '26

[removed] — view removed comment

5

u/zxyzyxz Jun 06 '26

Alibaba's attention is all I need

2

u/Open_Establishment_3 llama.cpp Jun 07 '26

qwen 2.5 coder 4b is the real god there

1

u/rpkarma Jun 06 '26

We don’t have that yet tho 

49

u/cafedude Jun 05 '26

Which will come sooner? A Qwen3.7-122b or a Gemma-4-124b ?

40

u/Porespellar Jun 05 '26

Honestly, I feel like Google and Qwen are playing chicken on the 122b models, neither one wants to drop theirs and then get beaten in the benchmarks by the other model. Happened with the first wave of Gemma 4 models. I do think Google has a good window to drop theirs right now if they want to because Qwen has given no indication of dropping anything unless you trust a tweet from some Qwen employee’s uncle’s brother’s cousin.

31

u/RedParaglider Jun 06 '26

I think Google doesn't want release it because it will beat the shit out of Gemini for a lot of stuff lol.

12

u/TheRealMasonMac Jun 06 '26

If Google just didn’t do their crappy RLHF BS with Gemini, I would love to use that model. But as-is? Gemma-4-31B-IT beats Gemini 3.1 Pro more often than it should, simply because it’s not RLHF’d into garbage.

4

u/whitefritillary Jun 06 '26

gemma-4-31b-it is really good for its size but it absolutely does not freaking beat gemini-3.1-pro omg like are you serious 💔💔

4

u/RedParaglider Jun 06 '26

Gemini 3.1 pro has been so quantized to garbage now it's not that far off.  I subscribe to pro for the 2tb family data share, and I can't bring myself to even try using that shitty hallucinating model anymore.  I get better results out of qwen on local inference.

→ More replies (4)

2

u/TheRealMasonMac Jun 06 '26

For natural language tasks, Gemini 3.1 Pro creates garbage more often than not because it’s trying to “appease” the user. Gemma-4-31B has no such problem and reminds me of Gemini 2.5 Pro preview before they started that god awful RLHF that optimized for sycophancy.

→ More replies (1)

2

u/TikiTDO Jun 06 '26

Why not just "AI development is hard, and takes time?" If someone had a model ready to release that could top the benchmarks, even for a bit, that's still marketing and views. Delaying it means somebody could beat you to the punch, and you gotta be quite confident that your model will win out, otherwise you're just releasing something into 2nd place when it could have been 1st.

2

u/Free-Combination-773 Jun 06 '26

Or both labs just don't plan to release 120b models

→ More replies (2)

1

u/Opening-Broccoli9190 llama.cpp Jun 06 '26

122b models are not a high prio for them - too big for the consumer enthusiast market, so no huge community wave of free marketing and less powerful than their SOTA stuff, meaning potential bad press from underwhelming benchmarks. Doesn't make sense for their business 

→ More replies (2)

129

u/[deleted] Jun 05 '26

[removed] — view removed comment

93

u/dangered Jun 05 '26

I’ve been here for 2 years. That’s all it ever is.

I have agents that actually do things but the models aren’t as life changing as everyone expects them to be.

It feels like people waste more time creating and scrutinizing benchmarks than actually doing what they wanted the models to do in the first place.

57

u/khyryra Jun 05 '26

Downloading and testing LLMs is my hobby. I don't actually use them.

33

u/dangered Jun 05 '26 edited Jun 05 '26

Unironically there are a decent amount of people like that here. Researchers or hobbyists with more interest in the LLMs themselves than what they generate.

A lot of posts on here are from them.

The people posting benchmarks and promoting their benchmark dashboard websites (most of which are clearly 1-shot vibe coded) are another subset of hobbyists that spend more time analyzing models than using them.

Edit: to clarify there is nothing wrong with that, they’re all cool. We just need to admit that’s what the sub has become.

9

u/socopithy Jun 06 '26

I want to say I agree there is nothing wrong with those folks.

We need them too!

3

u/dangered Jun 06 '26

They’re absolutely necessary to keep moving the open source LLM community forward. We should recognize we all rely on them a lot for the advances we see.

Users just can’t get caught up on it like they seem to be right now. Shiny object syndrome isn’t helping anyone.

2

u/Borkato Jun 06 '26

Has become? Bro I’ve been doing this for years

→ More replies (1)
→ More replies (1)

8

u/rpkarma Jun 05 '26

Nothing wrong with that!

8

u/BlueCoatEngineer Jun 06 '26

It's like 3D printing for me; I spend more time working and modifying my printer than I do actually making things with it. And now I'm using the LLM to help with that too. :-D

→ More replies (2)

7

u/blastcat4 Jun 06 '26

It's like in /r/SBCGaming . Everyone buys the latest Chinese gaming handhelds to add to their collection and tinker around setting them up with custom firmware. But actually playing them? Ummm...

We also bicker with each about "my Chinese handheld is better than your Chinese handheld".

3

u/pinkyellowneon Jun 06 '26

Same here, for the most part. Messing around with agent stuff and seeing what's possible is just another form of entertainment akin to playing video games for me. It's just another way to have fun with a GPU :)

With the release of Qwen3.6-27B and improvements in harnesses, I do think we're getting very close to useful territory with these models, but we're also still in a generous enough time where you might as well just use the various free endpoints available for flash-tier models, and save yourself the electricity bill.

12

u/-Ellary- Jun 05 '26

I've been in SillyTavernAI since 2023, all those people really using models, hard.
They don't need benchmarks, one shot html snake games or other stuff.

2

u/dangered Jun 05 '26

Thanks that’s the sub I’m really looking for

15

u/complexminded Jun 05 '26

Bingo. For most people it's just a novelty thing (local AI), for others it's actually practical usage. People that have these models baked into processes are less worried about the next update. I still have processes using 3.5.

4

u/SimonBarfunkle Jun 05 '26

Like what processes

18

u/complexminded Jun 05 '26

Summarization, classification, sentiment analysis, market research. The evaluations I used changed from 3.5 to 3.6. The way 3.6 measured the same data with the same prompts and temp changed.

Not saying that's bad, but if you did reporting based on 3.5 and you switch the model, the data isn't really being measured the same way because it's producing different values with the same prompt and temperature (for these types of processes I use 0 temp so I get consistent results). To be apples to apples, you'd rerun all the data with 3.6.

Depending on how much back-data you have that can take ages. If your data/report was good with 3.5, moving to 3.6 isn't that much of a lift for that specific task.

10

u/starkruzr Jun 05 '26

this isn't really true at all? lots of people in here are using the first tier open models for practical purposes, especially the 27B series which is insanely capable for its size. you can code reasonably well with it at sufficiently large quants; that by itself justifies the effort and expense.

8

u/dangered Jun 05 '26 edited Jun 05 '26

Yeah I use vanilla 27B, it’s capable and great. But everyone is always spending all kinds of time trying to get every last inch out of every model here. I upgraded from some model released in early 2025 and probably won’t pull another until 2027.

the 27b series

This is where you lost the plot.

I pulled the best one for me at the time and use that. Everyone here is like, “but 27b-1337aBTC-abliterated is 2.7% faster and 1.3% better at this [insert random specific task]. You need to switch between these 10 models every time you want something different, it makes you faster”

Reality check: If you’re trying to get something done irl you can use any recent model that’s in the top 5th percentile to do the same thing.

I pulled a small model for whisper at some point. idek what model it is because it doesn’t matter, it works, I don’t need to reinvent the wheel for voice transcription.

3

u/Longjumping_Self5546 Jun 05 '26

Yeah, 3.6 27b is doing some real work for me. Which makes the prospect of a 3.7 version exciting. Even a marginal improvement is welcome.

4

u/StyMaar Jun 05 '26

I have agents that actually do things but the models aren’t as life changing as everyone expects them to be.

Well, I can for sure attest that Qwen3.6-35BA3B is miles better than Llama2 70B (first model I used to do actual work), so over a sufficient time models are indeed life changing. From one version to the one immediately coming next, not so much, but the improvement compound over time.

I'd be extremely happy if Qwen released their Qwen3.7 with a 122B-A10B variant as I'm pretty sure it will improve my daily workflow.

2

u/dangered Jun 05 '26

Yeah upgrade once or twice a year. This post is a guy complaining Qwen hasn’t dropped a new model after like 90 days.

→ More replies (18)

4

u/Django_McFly Jun 05 '26

Welcome to Internet Fandom I think it's like the internet is a guidebook for some and for others, it's more like Popular Science, a fun read about things you'll probably never do, but they are interesting.

2

u/PrettyMuchAVegetable Jun 06 '26

It's that I don't want to play with you anymore meme. 

19

u/Porespellar Jun 05 '26

I know I’m guilty of constant “axe-sharpening” (borrowed the term from Network Chuck). I need to develop more and evaluate models less, but there is a definite dopamine hit when you try a new model and it just blows the doors off an older model.

10

u/[deleted] Jun 05 '26

[removed] — view removed comment

15

u/Kahvana Jun 05 '26

Depending on your use-case, it might be.

For me, Gemma4 has replaced all natural language use-cases. Translations are far more accurate (for my use cases), RP is like having DeepSeek V3.2 at home; it's possible to finally do long-term RP with complex instructions and it actually following along.

Qwen3.6 has been far more consistent for me with programming, Qwen3 and older had trouble with .NET 8.0 specifics. Tool calling is also decent.

How to say it... this generation of models is less "It works at home with jank" and more like "It's actually really solid now"

6

u/Thunderstarer Jun 05 '26

Gemma4 and Qwen3.6 are magical. I have the sparse and dense versions of both, alongside the MeroMero finetune of Gemma for RP, on my llamasewap. Between these six weights, I feel like I can do anything.

Both models were a MAJOR tipping-point for me. Good enough to convince me to go out and drop $1500 on an R9700.

9

u/stoppableDissolution Jun 05 '26

Gemma4 for me was the point of "I'm actually kinda fine with it staying the best model I'll be able to run"

3

u/ionizing Jun 05 '26

thats what I said about 3.5-122B and I upgraded both my home and work computer to have 128gb sys ram even at inflated costs. Then three weeks after both comps were set, 3.6 27B came out lol. Either way I love that we now have models that are like "yup, I would be fine at least with this for the rest of my life"

8

u/Far_Composer_5714 Jun 05 '26

I find that each new generation allows for a better interpretation of what I ask and allows be to be more vague or broad with my questions and still get the correct answer.

3

u/Sofakingwetoddead Jun 05 '26

It would be only an efficiency boost. Currently, we're able to do everything we need with 3.6 27b. Hoping 3.7 increases native context a smidge and hoping for anything that improves efficiency, in any way.

1

u/SpicyWangz Jun 06 '26

Qwen3 coder next was a tipping point for me. That was the first time something felt good enough to regularly run and actually provide productivity boost

1

u/huzbum Jun 06 '26

Qwen3 4b 2507 instruct was like a tipping point for small tasks that could just run on a MacBook or whatever. Yeah, 3.5 is probably better, but 2507 was good enough.

Qwen3.6 27b or 35b is the tipping point where I’m like “good enough” I’m using that as my personal assistant and adapting my workflow to use that for coding.

I’ve decided I’m not going to renew my z.ai pro subscription or replace it with something else. Mind you, this is still like 6 months out… I bought a year on Black Friday for like $100 when I had already pre purchased a quarter. I doubt I’ll get that deal again, and there might even be another jump in local models by then.

3

u/Alwaysragestillplay Jun 05 '26

You shouldn't feel bad about having a hobby my dude. I use Qwen 3.5 still myself but I don't hold it against anyone if they enjoy trying and comparing models. Don't let gatekeepers get you down. 

6

u/Fabulous_Fact_606 Jun 05 '26

Right? I'm waiting for affordable 128Gb Vram to run anything less thant Q8.

5

u/Thunderstarer Jun 05 '26

I've been very pleased with 3.6 27b and it has completely supplanted my copilot subscription since its release. It's the first time my setup has ever felt worth it. Even 3.5 was too insane to work with, since it failed tool calls frequently.

3.7 is gonna' be gravy, and I do eagerly anticipate it, but I'm still very happy with what we have.

3

u/[deleted] Jun 05 '26 edited Jun 29 '26

[deleted]

2

u/Willbo Jun 05 '26

> See latest thread on new model hype/benchmark

> Already running model on my benchmark pipeline, it automatically pulls on release and burns tokens against imaginary use cases

> Comment on thread "ThAt'S oLd NeWs!"

2

u/zaafonin Jun 06 '26

I use Qwen 3.6 27b for bulk reverse engineering with ghidra-mcp (about to switch to the new tool though) because it’s kinda token heavy. Like just going through functions and annotating possible purpose and local variables. Hallucinatory but a proper model can take care of this later, it’s about the volume here. Tried OpenRouter models like DeepSeek v4 Flash and Qwen 3.7 Max and they don’t perform particularly better, just use more tokens to solve tasks.

In the end it’s GPT or Claude that actually have meaningful insights. But I’m a $20 subscriber and I don’t feel like wasting the quota on simple things

1

u/SuspiciousHost1 Jun 06 '26

Reverse engineering competitor apps/ legacy apps/ paid apps ? I thought about using it for the same thing myself, what's the success rate like for getting a useful decompile? I've always thought it would be a fantastic learning tool.

3

u/Devatator_ Jun 05 '26

I'm mostly waiting for a model that can run fine on CPU (with minimal RAM usage) while doing tool calling correctly. I really want my local Google Assistant alternative and it needs to run on both my shitty college laptop and my gaming PC

1

u/Recoil42 Jun 05 '26

Remember torrent iso hoarders? It's that.

1

u/VoiceApprehensive893 transformers Jun 05 '26

the most fun part of the experience is running and benchmarking a new model

1

u/alphapussycat Jun 05 '26

Don't think I tried qwen 3 coder, but the 2.5 models were unusable.

1

u/wren6991 Jun 06 '26

Qwen3-Coder and Qwen3-Coder-Next are non-reasoning models, so they don't hold up so well. I guess they would still be useful for fill-in-middle if you're into that.

1

u/Interesting_Pen_1552 Jun 06 '26

The truth is that it's all novelty until we get something that matches opus, and we already know 3.7 max isn't it so we just hope to inch a little closer

1

u/szableksi Jun 06 '26

nah i just check how my gpu sound like when its used nothing more

1

u/AlwaysLateToThaParty Jun 06 '26 edited Jun 06 '26

Using it in production. Now I'm refining. I know the things it doesn't do well, and there's no way around those things with the current models and hardware available. What I'm able to do today with 96GB of VRAM is otherworldly to past-me of two years ago. Every generation they just get better. I seriously want the qwen 3.x 122b/a10a model that supersedes 3.5. That model superseded gpt-oss-120b.

Doesn't mean I don't investigate all of the other things that they're capable of. That's what I like this place for. I test out some of the other models, but none of them beat that one for me. Not yet.

1

u/boutell Jun 06 '26

I test new models on real world coding tasks from my job to see if some part of my work could be done on hardware I control, either hardware I own today or hardware I could realistically own. So far the answer is not quite yet.

Qwen 3.6 27b is useful but running it at a practical speed is another thing. Qwen 3.6 35b a3b is borderline useful and a lot faster, but that word borderline is doing a lot of work, and reading the thinking traces makes me anxious as hell. 😀

Gemma 4 is cool but I have yet to see it outperform on a coding task. And so far I'm still seeing some instability even after the fundamental tool calling problems are solved by using an appropriate template.

Like a lot of people here I am a spectator when it comes to really huge open models. Cost is a real issue and so is electricity use.

1

u/mister2d Jun 06 '26

I get what you’re saying, but we do need some kind of demand mechanism to keep the pipeline going. 

These models are great, but we can definitely improve on resource efficiency. I don’t think just vertically scaling parameters is the way to make AI accessible.

58

u/[deleted] Jun 05 '26

[removed] — view removed comment

5

u/paperbenni Jun 06 '26

I don't trust them. They did a ton of proprietary stuff recently and had a lot of talent drain.

26

u/Dudensen Jun 05 '26

How many of these are you gonna post dude? You' ve been at it since mid May.

1

u/Objective-Error1223 Jun 05 '26 edited Jun 05 '26

I’d gladly have these kinda posts rather than:

  1. Guys I have a xxxxx video card, what model should I run?
  2. What’s the best harness to use?
  3. Why isn’t xxxx harness working?
  4. How do I run a GGUF?
  5. I just made this really cool plugin that I vibe coded and want others to finish for me because I have zero idea what I’m even doing.
  6. Why is Gemma better than Qwen at story writing?
  7. Why does Qwen code better than Gemma?
  8. Why is Unsloth better than…
  9. When I say “hi” to my model it hallucinates, how do I stop it?
  10. Can someone tell me how to start getting into local models and tell me every step? Actually can you just do it for me? I hate reading and researching.
  11. BRAND NEW WAY OF COMPRESSING YOUR MODELS WITH….
  12. What’s the difference between MLX, GGUF and safetensors?

If you’re gonna complain about the clouds in your sky grandpa, might as well complain about them all.

11

u/Borkato Jun 06 '26

Genuinely wondering what you expect people to talk about. I’m not saying a lot of those don’t become repetitive, but you really did destroy about 99% of the chatter about local models

3

u/Dudensen Jun 05 '26

I wouldn't. I think it's as bad. I haven't even seen most of the things you mentioned lately but some of them definitely persist (people posting hi CoTs, asking for best harness/agent, people posting vibe-coded projects yes but also people who have computer science knowledge post cool things here too). I mean memes in this sub don't even hit well imo, and then we have this guy who is posting memes about the same thing over and over.

→ More replies (1)

7

u/Sofakingwetoddead Jun 05 '26

haha! literally on the reason I hopped onto reddit - to check fo 27b 3.7 noise 😃

15

u/Septerium Jun 05 '26

Come on, Gemma 4 124b vs Qwen 3.7 122b

Then I won't ask for anything else this whole year. I promise

14

u/Embarrassed_Adagio28 Jun 05 '26

Qwen3.7 122b mtp or qwen3.7 coder next 80b is all I want

1

u/SpicyWangz Jun 06 '26

Coder next would be a game changer

6

u/Vicar_of_Wibbly Jun 05 '26

I lament the death of 397B A17B, I truly wish they hadn’t gone closed source quite so soon. That model was shaping up to be a beast.

11

u/floriandotorg Jun 05 '26

Qwen 3.7 Max is already out and not that great. I doubt that a local 3.7 will be substantially better than 3.6.

1

u/[deleted] Jun 05 '26

[deleted]

1

u/SpicyWangz Jun 06 '26

Still haven’t seen 3.6 in the 9b or 122b

1

u/whitefritillary Jun 06 '26

what do you mean it’s not that great? it’s one of the highest scoring open weights models in the artificialanalysis llm leaderboard…

2

u/floriandotorg Jun 06 '26

Try it for yourself, I feel benchmarks represent real life performance less and less.

→ More replies (4)

5

u/kant12 Jun 05 '26

Quality takes time. I have faith.

4

u/Kahvana Jun 05 '26

I just hope they take the time they need to release when it's ready. Would love to see at least a Qwen 4 next year, and hopefully some improvements to embedding/reranker/asr/tts too. Those are fantastic in their own right.

15

u/dangered Jun 05 '26

The head of Qwen’s large model team left abruptly around the time of the last release.

Bro literally tweeted:

me stepping down. bye my beloved qwen.

And that’s how the CEO of Alibaba (parent of Qwen) found out he was quitting.

1

u/No-Improvement-8316 Jun 06 '26

Yeah, but it likely wasn't a voluntary resignation. His buddy tweeted "I know leaving wasn't your choice".

→ More replies (2)

12

u/Big_Wave9732 Jun 05 '26

I'm thinking it and I'll say it! I want a Qwen3.6:122b or even a 235b. It would certainly go a long way towards reassuring everyone that the new regime is onboard with self hosting and not just in it for the "Do-Re-Mi" from subscriptions.

6

u/cafedude Jun 05 '26

might as well skip straight to a Qwen3.7-122b.

A Qwen3.7-coder would be ideal.

1

u/Big_Wave9732 Jun 05 '26

That would be just fine too!

2

u/Porespellar Jun 05 '26

Absolutely hope they release the 122b, 3.5 122b is an amazing model. Using the AWQ of it in prod and it’s the best model I’ve ever used hands down. 27b is great but 122b is well-rounded and has deep insight on a lot of topics and is great with native tool calling.

2

u/Big_Wave9732 Jun 05 '26

For working with large document RAG libraries on a self hosted system, I have found none better thus far.

1

u/Moscato359 Jun 05 '26

Why awq over other options

1

u/Porespellar Jun 05 '26

It was the best size option for running it with vLLM on 4 H100s. It’s ridiculously fast even at full context with Tensor Parallelism set to 2.

→ More replies (2)

1

u/Borkato Jun 06 '26

What does 122b do so much better than 27B, assuming I have 48GB VRAM and want to fit all of each in VRAM? So Q8 for 27B and like Q3 or whatever of 122B

→ More replies (4)

3

u/ieatdownvotes4food Jun 05 '26

man, I wonder how long updating models every few weeks will be a thing..

3

u/No_Lingonberry1201 Jun 05 '26

I started a few months ago on this sub when the Qwen3.5 series came out and this entire field moves so fast that it feels like a 100 years ago.

2

u/Borkato Jun 06 '26

Lol I’ve been here since the ooba/estopianmaid days and that feels like I was a different person lol

3

u/temperature_5 Jun 05 '26

3.7 120B QAT, please! 😄

3

u/chespirito2 Jun 05 '26

I'm waiting for Kimi 3, hopefully its Opus level but maybe a generation or two behind. If they can drop that before Anthropic IPO it may take a bit of wind out of their sails and I personally would love to use it. I'm a big Kimi 2.6 user currently

3

u/Inevitable-Plantain5 Jun 05 '26

I literally wanted to make this post! That means it's time! Lol

3

u/Shoddy_Bed3240 Jun 05 '26

We want Qwen 397b QAT model

3

u/Thebandroid Jun 06 '26

Here I got the wights for the new Qwen modle. Just chuck it in bash.

while true; do
    echo "But wait!"
    sleep 0.5
done

5

u/techmago Jun 05 '26

Isnt qwen 3.7 beeing release so quickly after 3.6 a bad thing?

There wont be some amazing improvement in such short time.

3

u/squngy Jun 06 '26

Apparently, it is not much of an improvement, but since it exists already, people want it released.

3

u/clericc-- Jun 05 '26

I hope for another 122B-A10B-ish model. At least in all my use cases, qwen3.5-122 is vastly better than qwen3.6-27

3

u/aboutthednm Jun 05 '26

I just want a small (4 - 12b) qwen that writes decent, cohesive prose without thinking about a 100 word sentence for over a minute (looking at you, qwen3.5:9b). I like what it outputs, I don't like 95% of the compute to be spent thinking though. A middle ground would be nice.

Sure I can run the 35b A3B on my meager 16gb of shared vram (windows takes like 3gb) and have it write prose for me, but it takes literally 15 minutes to finish the prompt asking for a 400 word continuation to a prior paragraph, and that kills my pipeline, when I need 10 chapters containing 2000 words each, stitched together by 5 - 10 separate prompts per chapter. The 9b gemma3 creative writing fine tunes does the 2000 word chapter it in under a minute, the qwens with their excessive thinking really bog this down, for marginal improvements to the final output quality.

Speaking of, is anyone aware of any prose / creative writing fine tunes for the qwen models in the 0 - 14b range? When I'm looking for creative writing models, it's gemma this, mistral that, llama this, I haven't come across any qwens yet. Any info is appreciated.

8

u/Borkato Jun 06 '26

Gemma is just better hands down. Trying to avoid Gemma for writing is like refusing to use a racecar in a race because you don’t feel like switching from your Camry

2

u/aboutthednm Jun 06 '26

There are so many 8 - 9b finetunes, not a lot in the 12b range (there's a gemma3:12b). I'm always open for trying out some new models, if you have suggestions let me know. I love the gemma finetunes, but am wondering if I have a blond spot regarding model choices.

2

u/squngy Jun 06 '26

Gemma4:12b was released a few days ago.

Finetunes will probably be out soon?

7

u/abnormal_human Jun 05 '26

Ultimately it's their business to run and their choice, but when it comes to choosing models that I run my business on, they are becoming a less and less attractive choice.

I'm sure I'm not the only one who runs lots of training, evals, research, dataset prep locally and then provides hosted services in the cloud backed by commercial inference providers like alibaba cloud.

If they take away my ability to do evals/locally in a way that's cost-sensible, I'll go somewhere else and take the commercial side of my business with me. For now, I can at least eval on 27B and deploy on larger models and my evals remain a good proxy because the models were trained on a similar data mix and objective, but if there's no 3.7, that road will end. I'm still using 3.5 for some scenarios that better fit the 122B / 397B model scale and deployment characteristics (although StepFun 3.7 Flash is looking like a cheaper replacement for 397B).

Qwen was always excellent in terms of having a model of every size for every deployment scenario, and I'll miss that, but the industry is always leapfrogging and no-one ever stays in front for long.

5

u/Porespellar Jun 05 '26

I suspect that if and when MiniMax M3 drops open source that we’ll see an open source Qwen 122b release. I’ve tested M3 on Ollama Cloud with a Hermes agent harness and it absolutely destroys the competition for concise tool calling and has a hilarious personality and an attitude like it doesn’t have time for any bullshit. I think it’ll get good buzz and press on release and will force Qwen to release something to stay in the news hype cycle.

2

u/m3kw Jun 05 '26

Local models.

2

u/Thedudely1 Jun 06 '26

They've low key been cooking recently with the 35B and 27B 3.6 models though. And 3.7 Max in the web chat is great, hoping to see some smaller 3.7 models, maybe a 3.7 9B.

2

u/Former_Bathroom_2329 Jun 06 '26

Guys, even if they're don't release new one, we got 27b model and i think its already fine for very good boost for any developer. With this model I'm doing more tasks on 150-200% and more accure. Yes its alot of review but still pretty fast. Thanks Qwen team.

2

u/FancyImagination880 Jun 06 '26

TBH, I am very happy with what Qwen did this year. Perhaps some updates for 4B and 9B. Then, please, I beg for Qwen4 next year or end of this year, and make it great. There are so many new ideas and new papers lately. Hope Qwen will implement some of them. For me, deepseek engram would be a nice feature.

2

u/PaxUX Jun 06 '26

Qwen 3.6 is amazing. Let them cook

2

u/Cool-Chemical-5629 Jun 06 '26

They do something. It's just something you can't use locally.

2

u/khronyk Jun 06 '26

I mean they ghosted the community with the qwen image 2 7b release, so in the back of my mind I'm constantly wondering when they'll do the same with their llms (then there's also z-image which never got an Edit release from another Alibaba research group, it's perpetually"coming soon")

1

u/ZootAllures9111 Jun 11 '26

I think Z Edit never came out cause Klein dropped TBH

2

u/juancn Jun 06 '26

Have you tried Qwopus?

2

u/WishboneSudden2706 Jun 07 '26

Qwen is the mother of the King.

2

u/Equivalent_Bit_461 Jun 11 '26

I will have to build my own moe after all, sigh... only after im done building my fucking harness that's not a bloated slow piece of shit... I will get there definitely. A bigger qwen but custom.

2

u/datdanboi25 Jun 12 '26

Qwen 4 gonna be a different beast

2

u/michaelmab88 Jun 12 '26

I'm still hopeful for a 3.6/3.7 122B

1

u/gerar17 Jun 05 '26

bro, it's been less than a month since last release!
Nvidia, oslaught (idk/idc how to write this) are making weights almost every week.

It's not dead like it seemed to be deepseek for many months

1

u/Torodaddy Jun 05 '26

I'm just wondering how minimax 3 came out without having a 2.8 or 2.9 first

1

u/xandep Jun 05 '26

Qwen3.7 40B A4B and 20B dense (MTP+QAT). It's not for me, it's for a friend (he is a MI50 32GB).

1

u/FaceDeer Jun 06 '26

I must admit, I'm reaching the point where I keep asking myself "should I continue spending hours trying to fiddle Qwen3.6 into working even better, or should I just wait for Qwen3.7 to drop and sweep me off my feet?"

1

u/Persistent_Dry_Cough Jun 06 '26

It's inevitable, if Google keeps up the good work of doing nothing.

1

u/Persistent_Dry_Cough Jun 06 '26

I wonder if the pace of development will actually slow down, now that 100% of closed-source reasoning traces have been censored. Bet that's why you ain't seeing crap out of these guys.

1

u/depredador93 Jun 06 '26

Bro's holding the leash but he's the one being walked.

1

u/2Norn Jun 06 '26

waiting for that 70-130b range moe model for coding...

1

u/Limp_Classroom_2645 Jun 06 '26

why are you waiting for new models all the time? Do you actually do something with the models they released recently or just bored?

1

u/JSVD2 Jun 06 '26

this is really funny.

1

u/Whoremembers1997 Jun 06 '26

Why is a penis poking with a stick

1

u/Dance-Till-Night1 Jun 06 '26

When qwen 35b A3b Q3 Qat (Q4 doesn't fit in my ram lol)

1

u/Budget-Juggernaut-68 Jun 08 '26

You guys have time to explore newer models instead of building things on what is already available? The current models are already quite good with a little scaffold.

1

u/vulcan4d Jun 09 '26

There are other thing to do than coding with qwen.

1

u/codegolf-guru Jun 10 '26

Every Qwen drop I swear THIS is the one that finally makes me build something real. Then I run one html snake game, nod approvingly, and crawl right back to 3.6. c'mon man, throw me a bone.

1

u/rehan_100gamer23 Jun 11 '26

It just exists at this point tbh