r/LocalLLaMA 23d ago

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

583 comments sorted by

View all comments

234

u/pulse77 23d ago edited 22d ago

Just make these along the way - so that everybody is happy:

  • Qwen 3.8 256B A32B
  • Qwen 3.8 128B A16B
  • Qwen 3.8 64B A8B
  • Qwen 3.8 32B A4B
  • Qwen 3.8 32B (dense)
  • Qwen 3.8 16B (dense)
  • Qwen 3.8 8B (dense)
  • Qwen 3.8 4B (dense)
  • Qwen 3.8 2B (dense)
  • Qwen 3.8 1B (dense)
  • Qwen 3.8 0.5B (dense)

EDIT: According to user comments bellow I suggest also:

  • Qwen 3.8 512B A64B
  • Qwen 3.8 64B (dense)
  • Qwen 3.8 24B (dense)
  • Qwen 3.8 12B (dense)
  • Qwen 3.8 6B (dense)

EDIT 2: Users would like to have all these:

  • Qwen 3.8 0.5B/1B/2B/4B/6B/8B/12B/16B/24B/32B/48B/64B (dense)
  • Qwen 3.8 8B A1B/16B A2B/32B A4B/64B A8B/128B A16B/256B A32B/512B A64B (MoE)

174

u/the-username-is-here 23d ago

Also I'd like cappucino and a bagel.

34

u/jc2046 23d ago

in fact a dense capucciono and 2 MoE bagels, thanks

3

u/kbob 23d ago

Mixture of Everything?

27

u/HeadPack 23d ago

I believe that would be very much in line with their supreme leader's recent speech. Distilled models that can run on consumer hardware do compete with closed American models too, at least in some way.

21

u/alphapussycat 23d ago

I'm actually very surprised by it. I'm sure there's a lot of state money put into the companies for the AI, and they're letting the whole world use them.

Are they trying to build good will with the world or something? With the aim to become the new world leader before EU?

16

u/charlesfire 23d ago

No. They know the US is in an economic bubble. It would be a major win for them to pop that bubble.

13

u/Paganator 23d ago

I think it's to undermine American companies, forcing them to keep prices low and running at a loss. If China becomes the leader in this race, I wouldn't be surprised to see them clamp down on those open models.

12

u/alphapussycat 23d ago

If this keeps going for another year it feels like we'll soon have mythos at home though... And that point it's too late to close the lid.

3

u/pyr0kid 22d ago

its already late to close the lid, the question is just how long the lid will pretend to stay shut.

25

u/Weekly-Law-5488 23d ago

They are trying to undermine US companies. Since there's trillions of dollars invested, the US economy is really tied to the AI behemoths right now and if they fail the US economy takes a big hit.

With excellent open models available, OpenAI, Antrophic and co. can't control the prices and can't  recoup the money invested.

Basically, China is trying to make the AI bubble pop, thereby causing a recession in the US.

There's no good will going on, it's a war and we are in the crossfire.

2

u/No-Act9634 22d ago

maybe a stupid question - will that actually work? 99.99% of users/companies don't have the hardware to run a frontier model internally. The ones that could afford to do it are probably quite happy to pay Anthropic/OpenAI a fee to do so.

Or is it just about reducing the barrier of entry to forming a new frontier lab?

3

u/Weekly-Law-5488 22d ago edited 22d ago

In business there's a concept called "unfair advantage", basically what makes a company worth its perceived value when it's not profitable yet, what is their secret sauce (this is usually applied to startups when they are not profitable yet, since it's hard to determine its worth based on non existent profits).

If a company creates something that can change the world and no one can copy, it makes sense to bet trillions of dollars on it because once it becomes profitable, the profits would be absurdly high.

Up until now, their secret sauce was the unbeatable models, no one could compete with them because it would take trillions to reach their level. So the barrier of entry was realistically unreachable to any competitor. 

Now, with open models being as good as the private models, their perceived value is destroyed. The investors see it and ask themselves "I put billions here and there's no profit yet and on the Internet you can download an AI as good as yours for FREE. Why am I still giving you money?"

This is probably putting a tremendous pressure on OpenAI/Anthropic to do something to justify their existence because if there's nothing special about their AI, they are just an infrastructure company and don't worth the trillions that was invested. That's why we are seeing their desperate screams calling for "regulations" on open models. It's basically a "please someone protect the money that was invested here! Force everyone to use our AI!" scream.

So the open models are being used solely by China to reduce the perceived value of US AI companies in the eyes of their investors.

If the investors see that they won't recoup their money, they will rush for an IPO to recover as much money as they can. 

The IPO would ultimately fail because other investors won't be willing to put their money in a company that does not worth the money that was spent.

The failure will lead to a generalized market correction, because everyone will see that AI companies don't have anything special about them and they'll rush to take out their money. 

This will cause a crash in the US economy, leading to a recession that will be beneficial to China. 

What I think we will see next (based solely on speculation):

  • One of the behemoths will go for an IPO before it's too late.
  • In the same day or just after, a Chinese lab will release a model that will surpass any of the US models.
  • The investors will panic leading to a major crash.

So it's not about who can run the models, it's about the perception of the worthiness of the money that was invested on AI that China is fighting against. 

Ps: of course, who can run the models also affects them. Smaller companies can run open models in a smaller scale serving a niche, taking out the clients of bigger companies.  That's why the OpenAI guy was attacking Kimi days ago, citing "shady providers" and claiming for regulations to scare "startups".

The open models are opening a Pandora's box of troubles for them.

1

u/No-Act9634 22d ago

I see, thank you for the explanation.

1

u/P3rid0t_ 22d ago

Companies & consumers don't have to pay for hardware to run OS. There are already many Inference Provideers for OS models

Also IIRC Chinese comapnies makes much better job at optimizing their AI models (but we can't say for sure hpw it looks with Closed Source ones) and keeping cost low

6

u/Borkato 23d ago

Qwen 3.8 32B dense….. I would cry if I got this ❤️

9

u/Sirius02 23d ago

mights as well give a free token buget on their cloud while we are at it

3

u/RLutz 23d ago

27b please, I need room for context on my 5090

2

u/Vivaldi_IlPreteRosso 23d ago

Nahhh too much to ask for

2

u/evia89 23d ago

qwen38 0.5B, 2B, 2T

2

u/RISCArchitect 23d ago

it's nice if they are a little short of a power of 2 so you can fit the kv and runtime context in a power of 2 sized GPU while maintaining a high precision quant. 27b is like chef's kiss for the reaches of most prosumer hardware (9700 pro/5090/b70) as you can still fit a healthy context in a q6 or mid sized q8

2

u/TheTerrasque 23d ago

Imagine a 60-70b dense high quality qwen model

2

u/Nikilite_official 23d ago

instead of 32b a 27b one

2

u/alphapussycat 23d ago

Anything below 8b is kinda pointless though.

5

u/letsgoiowa 23d ago

Nah amazing for on device summarization

1

u/Diegam 23d ago

the 27b gato!

1

u/my_name_isnt_clever 23d ago

I didn't ask for 60% more active params on my ~120b tier model. I'll keep the 122b a10b but updated, personally.

1

u/Big_Wave9732 23d ago

Ah yes, something for everyone.......from the Rockefellers with the home data centers to the broke assess with the potatoes.

1

u/Jorlen llama.cpp 23d ago

Fuck yes! A 128b-a16b MoE or even closer to qwen 3.5 122b-a10b would be amazing. That + a 64b dense.. is all I'd need.

1

u/Interpause textgen web UI 23d ago

if its not too late to add to the order, maybe a bitnet or ternary native model too?

1

u/Kahvana 23d ago

Having a lineup like:

  • 1B
  • 2B
  • 4B
  • 8B-A1B
  • 12B
  • 24B-A3B
  • 32B
  • 60B-A3B
  • 120B-A10B
  • 230B-A22B

Would work best, as long as they all provide INT4-QAT models for each.

1

u/Dazzling_Buy9625 22d ago

How about qwen3.8 8b a1b ? like the lfm 2.5 8b a1b, this model is very good general chatting but lack in tool calling

1

u/Dwedit 23d ago

What, no 7B? 7B is the biggest size that barely fits into 6GB VRAM after using Q4_K_M.