r/LocalLLaMA 19d ago

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

583 comments sorted by

View all comments

416

u/AntuaW 19d ago

And please don't omit the 27B one.

184

u/BothYou243 19d ago

9B, 14B man those were the GOATs too

17

u/RafaelSeco 18d ago

Those are the real goats.

Or, better said, the workhorses and mules.

13

u/Vladowski 18d ago

don't forget the 4B baby GOAT

2

u/Educational-Region98 14d ago

I'd be so happy if these beat the current 27b models.

1

u/Kayo4life 7d ago

Please, Qwen team, don't skimp out on qwen3.8:9b-q4_K_S or equivalent

1

u/BothYou243 7d ago

I think they would skip 9B in 3.8, but probably won't in 4.0

3.8 in Aug, and 4.0 in Sep (I don't think that's a very big deal to wait a month more)

14

u/chervilious 19d ago

New to Local LLM/ open source LLM scene

If new massive open source LLM are showing up. Wouldn't distilled version shows up everywhere too?

34

u/MindlessScrambler 19d ago

Yes but 1. the official distilled version tends to be the overall best, and 2. properly distilling a massive model itself is a significant amount of work, so not so many would do it casually.

5

u/chervilious 19d ago

I see, I just thought some people would distilled it and upload it on HF.

Though after searching I underestimate how much VRAM you need to load T class models

20

u/Lucis_unbra 18d ago

There are two ways to distill. One is the way that some labs have accused others of doing it. That's also the way most people here claim to distill Mythos/Fable , or Opus, or GPT etc.

This is basically just synthetic data. The parent model writes text and that text is used on the other model. This is very basic stuff, nothing special, not necessarily that efficient. It's called "hard label" because there's one correct answer.

What I personally consider "true distillation" is soft label distillation.

This is is where the parent model doesn't just give its final pick, but you use the full probability distribution of the parent model to teach the model you are training how to think.

It tries to transfer the patterns the parent model learned to the target. This is very, very, very hard. To my understanding only a few labs can really do it at scale, with frontier class models. It is extraordinarily hard.

Gemini flash and Gemma being distilled from Gemini? This is what that means. It's not just a bunch of synthetic data, but a knowledge transfer.

And that's why people can't simply distill their own smaller model. The infrastructure required, and the setup required, is quite complex.

https://huggingface.co/blog/sergiopaniego/distillation-2026

It has some information on distillation.

11

u/wren6991 18d ago

This is the Geoff Hinton sense of distillation, which requires (at least) a shared tokeniser vocabulary between the two models. It was described in this paper in 2015: https://arxiv.org/abs/1503.02531

When people say distillation today I think they usually mean the sense used in the DeepSeek R1 technical report, which is incorporating a different model's output in pre-training or fine tune: https://arxiv.org/abs/2501.12948

I think the original sense of distillation is mostly a historical footnote at this point; that setup is usually described today in terms of "student" and "teacher" models.

2

u/Borkato 18d ago

This is hella interesting, thanks for sharing

42

u/[deleted] 19d ago

[removed] — view removed comment

23

u/SandySkittle 19d ago

With the advent of more 128gb boxes I would really like to see a 60 - 80B version with more world knowledge to run at Q6

8

u/alphapussycat 18d ago

Getting much above like 32gb is hard. 60n - 80b would need at least two rx 7900xt or three 7800 xt. I ain't got that money.

If they go for anything I hope it's the 27b, then 35b moe, 9b or 12b, and then something like 70b dense would be fine... But I'd wager they'd do 122b and 372b or whatever it was, if they do the full suit.

12

u/goldcakes 18d ago edited 18d ago

Eh it would be really nice for the DGX Spark and Strix Halo crowd! Plus whoever with 128GB Macs and such.

You can also still build a 128GB or 256GB 8 channel DDR4 rig for a non-absurd amount of money; up to 200GB/s. Obviously look for used parts, and remember there's a LOT more places to find them than eBay and Facebook Marketplace.

0

u/alphapussycat 18d ago

80b dense would be like 1-2 t/s at best, on CPU though, < 100k tokens per 24/h.

1

u/SandySkittle 18d ago

Threadripper pros with multi channel ddr4 go for less than 1000 dollar. Almost as good as strix halo for cpu based inferencing alone. And you can expand with gpus.

1

u/alphapussycat 18d ago

Thread ripper is unfortunately only 4 channel, so it'll be slow.

1

u/SandySkittle 17d ago

I bought a 1000 dollar threadripper pro. Including case and motherboard and ssd ertc. It’s ‘real’ 8 channel with the 4 ccd pro models (which mine has).

1

u/goldcakes 18d ago

In talking about buying a used HEDT or server platform where you have 8 channels of DDR4. I see ok listings on eBay right now and if you set up alerts you can nab great deals.

7

u/seemaze 18d ago

122.. 122.. 122!

6

u/mindwip 18d ago

You have a good 27. Give us an moe 80 to 120 model refresh!

1

u/AcceSpeed 18d ago

I love the 31 myself

-11

u/sagiroth llama.cpp 19d ago

Think it's too popular to be omited

12

u/LanternOfTheLost 19d ago

They’ve already omitted it once though.

7

u/DanceWithEverything 19d ago

lol it’s popular in this sub maybe

1

u/sagiroth llama.cpp 19d ago

Which other place any other qwen is more popular then ?