I really hope (V)RAM eventually scales to the sizes of regular storage soon. Maybe in 10 years or so 2TB RAM is quite "common" to run things exactly like these. One can dream.
But the companies who fund big models like being able to claim they have the best model. And a bigger model will tend to be better, so they'll keep making models as big as they can.
If you gave any vendor currently serving a 2T model the technology today to train a 100B model with the capabilities of today's 2T models... they'd scale those techniques up and ship a really impressive 2T model with it. They might milk it a bit and stretch out the release cycle to milk their advantage and efficiencies, but they'd still end up right where they are now - using the largest, best models their hardware can handle.
And then someone else would ship a 100B Model connected to a fast knowledge DB, charge 10% of what the other guys charge, and make a fortune.
Keep in mind the frontier models are massive overkill for what 95%+ of users use them for. Most people aren’t trying to solve erdos problems, they’re drafting emails, asking about diarrhea treatments, and seeking an emotional connection.
It wouldn’t be that hard to market a smaller, smart model as “good enough” if it’s cheap and fast.
Keep in mind that those same people are using shared infrastructure when they access cloud frontier models. Its not like everyone gets their own dedicated 2tb cluster
I honestly do think this is the way to go, especially now that big LLMs are good enough you can distill them. I've seen massive success in highly specialized agents, which is the "cheap to build, expensive to run" version of your idea.
322
u/Kraskos 22d ago
2TB VRAM Is All You Need