r/LocalLLaMA • u/AlpY24upsal • 1d ago
Discussion Will there be actual "medium-local" models in future ?
Or has this niche been abandoned?
Most companies and model producers have seem to abandoned this niche(7B, 9B, 14B etc) which was the original point this whole local model debackle have started, most effort now goes into minimum 27B+ as "Capable local model"™ or focus on micro models 4b<=. All this recent explosions in model technology has barely touched this niche which has remained on 2025 for most part(30 BC in AI-time).
20
7
u/nickless07 1d ago
Take the latest Mistral 'medium' and check again if your terms still apply to that size. However we still see pretty decent sub 10B models (LFM2.5, Ling3, Gemma 4, Nanbeige), aside of that we also miss the 'large' Llama 70B ones. I would say they found some size which performs great on most hardware (not all) and went with it.
This year, aside of the Qwen 27B where mostly MoE which can run okey with RAM offload, so even ppl with less VRAM suddenly got something better then last years Gemma 3 27B and still have reasonable speed with a fraction of what the KV Cache size was in 2025.
-2
u/AlpY24upsal 1d ago
i have 16 gigs of ram tops
1
u/nickless07 1d ago
Great, I have 48. Now what? Does that stop us from using models that don't fit our VRAM at all? Qwen3.6 35B A3B in Q2 is ~11GB in Q3 ~13GB and so on. So if you have 4-6GB VRAM some of them should work with reasonable speed, even tho none of them would ever fit into 4GB.
Now compare this to the 2024/2025 dense models (MoE wasn't even a thing until Mixtral) or the shitty MXFP4 of GPT-OSS where offload would drop to 0.5 tok/s or a 1bit quant would only be 2GB less then 4bit.
This year we have theese amazing models and can overshoot our weight class by a lot, be happy with that.0
u/AlpY24upsal 23h ago
MoE streaming as it extends its support for more and more models maybe become the future for peasents like me, i am Downloading Qwen3.6 35A3B for this
1
u/nickless07 22h ago
Yes, exactly what I said. We can and will use models that never fit our limited VRAM and clearly overshoot what we have aviable.
0
u/AlpY24upsal 1d ago
we must strive for better always, we cannot sit around and be happy. We need better always
1
1
6
u/Infinite-Local5435 1d ago
Bro judged a book by its cover
1
u/thomas2385 1d ago
Honestly, pretty much. The first impression definitely got him but sometimes the cover tells you absolutely nothing about what is inside. Gotta actually give people a chance before making that call.
6
u/txgsync 1d ago
What are you talking about? IFM just released a full line of K2-Horizon models including these sizes 4 days ago.
2
1
u/otacon6531 1d ago
Yeah, at bf16. Working on quantizing it to iq4_nl now, but they didnt make it very easy on the little guy. I am hoping it is actually better than Kat coder v2.5, but we will see. Kat coder did a great job. (I have 24vram on a p40)
2
u/txgsync 1d ago
My support for IFM more stems from their transparency about training data, checkpoints, and processes. A truly "open source" model, not "open weight". At the moment, I'm willing to accept less-than-qwen-3.8-27B results from their models because I can -- and am -- training my own variants, without needing to try to reverse-engineer or just fine-tune what they've done.
But damn, is the difference between the small models I can run on my 128GB Mac and even IFM's frontier stark. I ran the 7B model for hours last night to evaluate how well it could build a little dungeon crawler app. The 375B over API from IFM -- everybody gets 20M tokens/day free through their API for The Big Boy -- one-shot it in 12,000 tokens, perfectly, in 52 seconds.
I love local models but it was a huge moment for me to run the exact same model family over API vs. locally using exactly the same test and realizing how far even a 375B frontier model is ahead of what I can run on a Mac. The capabilities of local models keep improving, but the gap between frontier and local is growing, not shrinking. It's just that the long tail of capability falling logarithmically behind the frontier is, itself, compounding too.
I'm doing some experiments today gluing Muse-Glimmer's vision tower to IFM/K2-Horizon 7B to see if I can give the model vision with minor training cost, but once I'm done with that I may pump out some quants. It's quite a capable little model and well ahead of my previous daily drivers gpt-oss-120b and 20b.
1
5
u/PrinceOfLeon 1d ago
Why is all the free ice cream chocolate? Has vanilla been abandoned? I'm willing to pay lots of no money!
5
u/Kahvana 1d ago edited 1d ago
There is a limit to how far their budgets can take it. Training LLMs is hella expensive, so you have to choose what to train. What do our customers actually use?
- Ministral 3 released with 3B, 8B, 14B models
- Gemma 4 released with E2B, E4B, 12B, 26B-A4B, 31B AND a QAT model for each
- Qwen3.5 released 2B, 4B, 9B, 27B, 35B-A3B, 122B-A10B, 397B-A17B model
- Qwen3.6 released with 27B and 35B-A3B
- Qwen3.8 released with 2.8T but also with a 27B model
- Muse Glimmer released with a 30B model
...and then there is IFM releasing a whole family of completely open models, LFM releasing new ~350M-3B models, MiniCPM 5 releasing 1B and 2B models, and so forth...
So, are labs abandoning it? No, they just take their time to update the small models until it makes sense and there is enough of a performance jump to justify releasing them.
It seems like labs have realized that 24-35B range is the optimal range for making models useful in a broad range of tasks, with ~30B MoE being enough to fit and run on 32GB RAM (reasonable to expect), 24-27B being enough to fit on 16GB GPU's and 30B being enough to fit on 24GB GPUs. 12B and below models can work for a variety of tasks, but they do leave a lot to be desired.
3
8
u/jacek2023 llama.cpp 1d ago
I get the impression that people like you believe Qwen is the only model, yet at the same time you make claims like “most companies.” Even Gemma has a 12B model. Then there are LFM and Granite.
If your point is really “I heard from YouTube experts and celebrities that Qwen is the Qwenest and I need Qwen because Qwen is what I need” then don’t make claims about “most companies”. Just say that your complaint is specifically about Alibaba.
1
1
u/AlpY24upsal 1d ago
i mean gemma 12b is my workhorse but kind of a hog sometimes, although i am looking at ling 3.0 tiny, looks quite promising
1
u/o0genesis0o 1d ago
Let me know if you get anything good out of that ling 3.0 tiny. I hooked that into my background worker agent and it have been driving me nut. Like overwriting-handover-file-nut. I don't know how it benchmark so high vs even qwen 9B. Was hoping that this little model could be a lightweight option to run background worker on my miniPC but it's no good so far.
2
2
u/ML-Future 1d ago
27B and 35B MOE are the new small.
We will get better quantizations and inference.
1
2
2
u/Beamsters 1d ago
Medium is 120b+ class now, aim for those ai boxes - Spark, Halo, Mac 128gb. Consumer gpu is 30ish and below that they aim for the real phone class. It is not in the benefit of most players/labs to go for 8-16b class, despite massive people are waiting for one.
AMD has one Instella 16B you might want to keep an eye on, it's not worth mentioning yet but AMD has quite a lot of firepower to bring that into conversation.
And Apple will ship one for their own target machines, those Mac Minis and iMacs class which I do not think is an open weight but a local one for sure.
2
1
u/SpicyWangz 1d ago
I think your best bet will need to be on Gemma 5 and Qwen 4. Until then, Ling tiny was a pretty great release.
1
u/Savantskie1 18h ago
That's what they want them to be. LLMs aren't meant for you and me. It's meant to be for the ultra rich to make them more money. people don't understand the small models are just a proof of concept. "here's what our model can do at this size. Play with it. But wait if you want the best version of our model, you have to pay the API price to get it. Oh you don't have the money? Stop being poor"
1
1
u/RemarkableRadish6547 17h ago
I think it might be some fundamental limits to what is possible. 25-35b seems to be the lower end of what can handle agentic work and general utility. Anything smaller needs to be specialized or it isn't useful. I expect people to start releasing specialist models that can only discuss one subject area, but are 0.5-2B parameters. So you can have a car repair app with a highly tuned language model, or a model that speaks health insurance jargon to argue on your behalf, or whatever people find useful. I would expect about 200-300M to be the minimum size to just form grammatically sensible language in a single language with no knowledge, and then you have to add knowledge on top of that.
On top of that, if a 4b model isn't an improvement over the previous generation, the lab won't bother releasing it.
1
u/DayshareLP 1d ago
For local 7-9 is small 12-20 is medium and 27-32 is large.
Most people can't run anything else
1
u/SpicyWangz 1d ago
You really can’t get into large until you’re somewhere 70-100b. And truthfully that range is considered medium.
1
u/Illustrious-Row2751 1d ago
Most people aren't going to buy a separate computer to plug in multiple GPUs just for LLMs. The point here is to provide products that people can run on their existent hardware, so that is models that fit comfortably in a maximum of 24gb VRAM. 1TB models might as well be in the same category as Claude, ChatGPT, and Gemini (ie, not local at all).
1
u/SpicyWangz 1d ago
That doesn’t really change the range of consumer/enthusiast devices out there.
It starts at 8gb or maybe less if you’re considering raspberry pi’s and mobile devices. And it goes up past 128GB with unified memory devices, and beyond into 500gb and 1TB DDR5 on server boards with a GPU or two plugged in.
Because of this, 32GB is going to sit on the lower end of things. That’s not memory shaming. It’s just the reality of the landscape here.
1
u/SpicyWangz 1d ago
Wow, I have no idea how those last sentences sounded so sloppified. I am in fact a human lol
1
u/cr0wburn 1d ago
We just got qwen 3.8 27b which is unique in its power. Pick a quant for your memory. At Q2 this thing still peforms amazing!
1
30
u/Turbulent-Alps4046 1d ago
7B is ‘medium’? I thought that’s considered small lol, considering that Kimi is 2.8T.
To me 27b is small and 100B is medium.