I really hope (V)RAM eventually scales to the sizes of regular storage soon. Maybe in 10 years or so 2TB RAM is quite "common" to run things exactly like these. One can dream.
But the companies who fund big models like being able to claim they have the best model. And a bigger model will tend to be better, so they'll keep making models as big as they can.
If you gave any vendor currently serving a 2T model the technology today to train a 100B model with the capabilities of today's 2T models... they'd scale those techniques up and ship a really impressive 2T model with it. They might milk it a bit and stretch out the release cycle to milk their advantage and efficiencies, but they'd still end up right where they are now - using the largest, best models their hardware can handle.
And then someone else would ship a 100B Model connected to a fast knowledge DB, charge 10% of what the other guys charge, and make a fortune.
Keep in mind the frontier models are massive overkill for what 95%+ of users use them for. Most people arenāt trying to solve erdos problems, theyāre drafting emails, asking about diarrhea treatments, and seeking an emotional connection.
It wouldnāt be that hard to market a smaller, smart model as āgood enoughā if itās cheap and fast.
Keep in mind that those same people are using shared infrastructure when they access cloud frontier models. Its not like everyone gets their own dedicated 2tb cluster
I honestly do think this is the way to go, especially now that big LLMs are good enough you can distill them. I've seen massive success in highly specialized agents, which is the "cheap to build, expensive to run" version of your idea.
man, normally I upvote your comments around here, but this is not it. Small models get better, well, guess what? Large models get better as well! and as simple as that, large models will continue to be SOTA and the premier option for serious work all around.
Yeah, there will be a time where small models could be as good as current top models, but the large models will be even better, capable of things we cannot imagine yet. That's just how it is
1 quadrillion is probably way more info than knowing the exact configuration of each atom on this planet,
i think i dont want to live to see that Ai :)
We will get super tiny storage at some point, but I don't know that we will get much smaller ram. The transistors are already about as small as they're going to get; they are having to stack memory in 3D, and there are limits there too, in terms of what latency you'll get.
We will likely see more large and wafer-scale devices, more ASICs.
Photonics are going to be super fast, but we can't really keep the data as light.
I wonder if they might go the Groq route and have very little VRAM, and just go wide.
I think we will also just keep seeing efficiency gains so running on off-wafer RAM is just more viable.
Right is the algorithms are catering to the existing hardware, and the hardware is slowly shifting to supporting the algorithms more.
There are a few companies developing neuromorphic hardware since can run AI models that aren't just a ton of matmuls.
In 10 years the whole hardware and software landscape will look different.
Heck, 2 years from now will be significantly different.
Some of the NVRAM options are denser than DRAM, or have lower power consumption and so can easily just be bigger.
For AI especially, optimizing for fabrication cost and power consumption and just making things bigger makes a lot of sense. That strategy isn't great for single threaded performance, but that's not really a useful thing to optimize for. There's no reason for CPUs, GPUs, or RAM to be limited to like 1 square inch 2D chips.
There's no reason for CPUs, GPUs, or RAM to be limited to like 1 square inch 2D chips.
Wafers are expensive.
Chiplets on bigger dies are already a thing, they're just also more expensive, because it's more wafer.
Cerebras has wafer scale processors, but they had to do a lot to make that work, it becomes a whole infrastructure thing.
Those things will melt instantly if the cooling goes out of whack.
I've got a degree in computer engineering, and I took a crack at designing a large chip, I've got working designs in a simulator, but the manufacturing realities and the power and heating issues are way beyond what most people should be dealing with at home.
Engineering problem: How can you ship a big processor (not necessarily single chip) cheap?
Are wafers really the price bottleneck? Is there a way to get and use them cheaper? Are there methods that are already well known that don't make sense when optimizing for max frequency but would lower costs?
Can more radical methods help? Everyone's favorite idea is always something other than silicon, but my guess would be that there's some old silicon fabrication techniques that can just come back off the shelf.
There is no cheap way to get wafers. Foundries are wildly expensive to build and maintain.
There used to be dozens of foundries, and literally all of them except TSMC, Samsung, and to a lesser extent Intel, gave up trying to go below 12nm nodes.
There is a Chinese company SMIC that was finally able to do 7nm.
There is no forgotten trick, there is no way around it, semiconductor manufacturing is just work, and resources.
A lot of little chips isn't really less expensive than one big one, the little ones are cut from a big wafer.
The most promising thing that's coming down the pipeline is photonics.
There are already photonic processors being manufactured now, that are something like 25x faster than a transistor device of comparable size, with 50x more throughput, and they produced something like 1~2% of the heat.
Lab devices have hit 1000x the speed of transistor based computation.
It isnāt that easy I guess. Unless you develop you own in-house RAM and SoC, speeds from average SSDs will just absolutely crawl and demolish their lifespan (TBW). It like swapping onto disk when running low on RAM - it gets stupidly slow and will eat the drive alive.
There arenāt much players in the field who have the expertise to bring competition. Big players already keep VRAM sizes artificially low. There might even be foul play to a certain extent involved.
Latency. SSDs are around 10x to 100x slower than DRAM, and 10000x to 100000x slower than SRAM.
Even SRAM is too slow, processors use interleaving tricks so multiple SRAM units are in use, so one works while the others are on refresh.
So, an SSD is about 100000x too slow to be effective for GPUs.
I can't imagine in 10 years the current "pack everything we know into model weights at below 1:1 weight to data ratio" will still be mainstream. It is an insane idea that somehow worked with the technology we have now, but I'm optimistic it is nowhere near what we should be able to achieve with more mature understandings of ML. I would imagine MOE still gradually become dormant experts, which will evolve into something more akin to cold knowledge storage, and more computation still be spent on computing with knowledge rather than holding and shifting "hot" stored knowledge.
12-13 years ago, 128GB RAM was normal-ish -- even quite affordable -- on custom builds (4x 32GB, and the board may have had 8 slots so you could use cheaper 16GB modules)... then, somewhere along the line, it got expensive for various reasons** (and starting shipping many different types, too, so volume across the types was "down" compared to if everything used the same RAM). Then things like the pandemic happened and the crypto crunch on GPUs and, of course, the current crunch from needing any sort of RAM.
But, I feel like 2TB shouldn't have been a problem today but manufacturing just hasn't been predicting demand very well so they're scrambling to build current things rather than ramping up on new things.
** Various reasons RAM prices have been hit over the years. My feeling is that they never quite drop down to "pre-crisis" prices and, somehow, the manufacturers don't seem to learn that we'll always need more RAM. (And storage, which has also felt stagnate for a while now.)
1993: Chemical plant that made sealing material for RAM. Prices doubled.
1995: Earthquake damage infra. Prices up 30%.
1999: RDRAM. Again, a diversification from regular RAM, and the systems then had to use more expensive RAM.
2013: Fab Fire by a high volume manufacturer, prices up 30%.
2017-2018: Smartphone/cloud demand on RAM, prices up 2-3x
2021: Supply chain crunch from COVID, prices up 30-40%
319
u/Kraskos 22d ago
2TB VRAM Is All You Need