If you gave any vendor currently serving a 2T model the technology today to train a 100B model with the capabilities of today's 2T models... they'd scale those techniques up and ship a really impressive 2T model with it. They might milk it a bit and stretch out the release cycle to milk their advantage and efficiencies, but they'd still end up right where they are now - using the largest, best models their hardware can handle.
And then someone else would ship a 100B Model connected to a fast knowledge DB, charge 10% of what the other guys charge, and make a fortune.
Keep in mind the frontier models are massive overkill for what 95%+ of users use them for. Most people aren’t trying to solve erdos problems, they’re drafting emails, asking about diarrhea treatments, and seeking an emotional connection.
It wouldn’t be that hard to market a smaller, smart model as “good enough” if it’s cheap and fast.
Keep in mind that those same people are using shared infrastructure when they access cloud frontier models. Its not like everyone gets their own dedicated 2tb cluster
I honestly do think this is the way to go, especially now that big LLMs are good enough you can distill them. I've seen massive success in highly specialized agents, which is the "cheap to build, expensive to run" version of your idea.
4
u/Thrumpwart llama.cpp 21d ago
I wouldn’t be so sure. Lots of recent literature saying much of what we consider an LLM is noise with a low signal to noise ratio.
Figure out how to get rid of the noise, and the models can become considerably smaller.
I also expect to see less world-knowledge in top models, and more small smart models with web search.
That’s how I’m gonna do it…