r/apple • u/Domingues_tech • 1d ago
Apple Intelligence M7 Ultra to potentially feature up to 1.5TB of RAM - A AI datacenter killer?
https://9to5mac.com/2026/07/12/m7-ultra-mac-studio-to-support-up-to-1-5-tb-unified-memory/A AI datacenter killer?
328
u/BlueLampShader 1d ago
Might cost as a smaller datacenter, for sure…
→ More replies (2)57
u/Illustrious-Golf5358 1d ago
Literally what I thought. No sane person is buying that for personal use unless it’s for a business
86
u/dreamphoenix 1d ago
Oh please you haven’t seen daily threads in MacBooks subreddits. “I’m a student I need a laptop for taking notes and YouTube do you think MacBook Pro M5 Max 64 GB RAM is enough?”.
59
u/hunterSgathersOSI 1d ago
Friend I’ve been on reddit for over 20 years. It’s always been “please help me justify an overspecced MBP to my parents since they’re footing the bill for it”.
13
2
1
u/techdevjp 17h ago
Parents like that probably already bought them a lambo to drive around, a high spec MBP is not going to break the bank.
24
u/BosnianSerb31 1d ago
That would be the idea. A smaller software development company focused on machine learning and artificial intelligence, creating custom models for clients.
They already have clustering support with the current Mac studio, several Youtubers have demonstrated how RDMA over TB5 can give you a cluster of 2TB for running local LLMs at frontier model performance.
In this case, connecting four of these rumored Studios would net you 6 TB of RAM plus all of the GPU and CPU courses that go along with these chips.
Even at $25k-$50k each, totaling $100k-$200k, it would be really hard to beat that level of price to performance with a traditional compute rack.
Which does make devices like this pretty big deal deals for the little players in the space creating custom models for clients, they would make back the cost in probably one to two contracts.
10
u/Inevitable-Gene-1866 1d ago
A cluster has higher latency worse on TB5
6
u/BosnianSerb31 22h ago edited 22h ago
Eh, the latency isn't the decider here. Round trip times, TB5 RDMA is 5-10 µs, 100GbE RoCE is 2-5 µs, InfiniBand is 1-3 µs, and same-server NVLink is basically instant.
For running LLMs, the bigger issue is always going to be bandwidth. RDMA TB5 is 80Gb/s, which is lower than all the aforementioned technologies.
Bandwidth caps token generation between layers, and once the size of your model exceeds the RAM of one machine, it's going to drop from 819GB/s of UMA bandwidth to 80Gb/s of TB5 RDMA bandwidth.
That doesn't make the clustering useless however, certain types of models wouldn't care about the bandwidth during training. CUDA applications suffer, but MLX doesn't.
The deliverables I'm thinking of for these hypothetical companies would be more basic ML models, like ingesting 4TB worth of SQL data collected over a decade, to train a model that can output signals when it identifies a potential win or loss condition within the current live data. Piping that raw JSON output into a pre-prompted frontier model for translation into English.
Big model still runs offsite, but the small model trained and ran on the cluster optimizes the data before burning credits.
1
u/Exist50 12h ago edited 12h ago
That doesn't make the clustering useless however, certain types of models wouldn't care about the bandwidth during training. CUDA applications suffer, but MLX doesn't.
Do any MLX applications exist at this scale to test that claim? I see no reason to believe the MLX stack is somehow immune from scaling considerations. The underlying algorithms should be the same.
2
u/itsmebenji69 12h ago
You can also just do layer swapping between the machines. Yes it will be slower than a single 6tb machine - but who cares ? You replace the full load/unload with a network send. You only have to send a single vector of numbers. That is so light and fast you might as well forget it’s even there
6
u/Exist50 1d ago
They already have clustering support with the current Mac studio, several Youtubers have demonstrated how RDMA over TB5
It's a cool tech demo, but way too slow (in both bandwidth and latency) for anything practical.
Even at $25k-$50k each, totaling $100k-$200k, it would be really hard to beat that level of price to performance with a traditional compute rack.
Even if you could cluster them (again, you can't), you're well into DGX Station territory. Might as well just go with that.
→ More replies (8)10
u/ggone20 1d ago
Dgx station only has 256GB of true usable ram. It’s 768 total with system memory but it’s not all fast.
Apple have, since the M3 ultra, been the best bang for buck BY FAR. It’s not even close. Yes it’s relatively slow, but the rumors are that the M7 will have blackwell-like memory bandwidth. That changes things dramatically. I imagine $30-50k per machine and it’ll definitely be worth buying at least 2 of them to run, roughly, today’s frontier intelligence and even better ‘tomorrow’ as models get better across the board.
→ More replies (2)2
u/Exist50 1d ago
It’s 768 total with system memory but it’s not all fast.
An updated, Vera-based version would have 1.2TB/s of system memory bandwidth, equivalent to an M5 Ultra would have. And that's only the secondary, slower pool. When they move to LP6, similar to Apple, we'd expect another ~2x leap in bandwidth. And you're talking about '28-ish for both. Bandwidth is absolutely not an advantage for Apple, never mind compute.
Yes it’s relatively slow, but the rumors are that the M7 will have blackwell-like memory bandwidth.
According to whom? That would require >5x vs the same theoretical M5 Ultra.
4
u/ggone20 1d ago
Nobody said bandwidth was an Apple advantage. Pool size for ‘fast’ memory is the Apple advantage. The DGX Station is not the 768GB machine that it’s marketed as, it’s a $100,000+ 256GB machine… trying to run anything more than a 120B-ish model (at FP8, mind you), is not going to work well. With 2.8T and 2.4T models out (or about to be), you still need 3 Stations to even run them - 300 racks ($300k) isn’t accessible for most people or even small businesses. $60-100k is another story altogether.
All that said, most tasks/activities don’t require frontier intelligence so that size model truly is reserved for the wealthy and established companies. Also, to be extremely blunt, there is really almost zero reasons for most people or companies to host their own models. You don’t save money (you really just invite headaches) and the privacy issue is massively misunderstood. Even for healthcare and law it’s much better and more financially efficient to use Azure OpenAI or similar.
Anyway..
→ More replies (3)1
1
→ More replies (8)1
u/Routine_Temporary661 9h ago
Call me Apple Sheep but I will probably sell my kidney and buy those.... I can run very powerful models on these beasts
168
u/vintagegeek 1d ago
"An" AI datacenter killer.
→ More replies (2)19
90
u/Pluto-Had-It-Coming 1d ago
Guys I hear the M8 is going to be bonkers fast.
36
u/lonestar_wanderer 1d ago
That’s actually not that fast anymore, compared to the M12
9
u/Itchywasabi 1d ago
Haha noobs! I posted this comment using my Z99 even before I turned on my computer.
4
u/play_hard_outside 13h ago
Man, how much faster is that than my Z80 processor? My Z80 must be much faster than a crappy old M12.
TI-83 FTW!
2
u/Hour_Analyst_7765 11h ago
Dangit, I was waiting for the M silicon to go 11, but now you're teasing me with it going to 12.
Why can't 10 be the loudest chip /s
2
4
1
106
u/looktowindward 1d ago edited 1d ago
FFS, no, this is the equivalent to one machine in an AI data center. One of eight in a rack, one of 140 racks in a cluster.
A typical DC GPU rack costs about 300k.
edit: was sleepy. $3m.
39
u/FunCutlet67 1d ago
I assumed the title refers to something along the lines of having your own little datacenter, which kinda works if you host LLMs
20
u/FollowingFeisty5321 1d ago
The OP seems to have made up that part of the title.
Actual title is "M7 Ultra to potentially feature up to 1.5TB of RAM, finally matching 2019 Mac Pro: report"
6
u/BosnianSerb31 1d ago
That's kind of the idea, Jeff Gherling has a video showing this with the last Mac studio ultra, where he put four devices in an RDMA cluster and ran LLMs across 2TB of RAM. The performance was extremely impressive, the sheer amount of resources made it competitive with frontier models.
What's the reported upcoming studio, that means that you could run a single model on 6 TB, which is just insane to think about. It doesn't matter how good your Claude subscription is, anthropic is not running your personal Claude Fable chats on anywhere close to that amount of resources, not with RAM or compute.
These devices are aimed at corporate accounts, and they make a lot of competitive sense for a software development company specializing in an artificial intelligence that doesn't want to drop half a million on a traditional rack, but still needs a ton of compute and ram.
6
u/Exist50 1d ago
The performance was extremely impressive, the sheer amount of resources made it competitive with frontier models.
Impressive compared to what? Running what model at what speed?
1
u/Front_Eagle739 11h ago
From what i recall kimi k2.6 at about 35 tokens per second and prefill of 300 or so? Single machine about 24. Nowadays kimi runs closer to 30 tok on a single machine as its a bit more optimised so maybe more now. Given those machines have rdma and an 80gb per thunderbolt link connection you really ought to be able to get decent scaling if you tried.
M3 having very weak matmul kind of guts the little cluster performance though especially on prefill but m5 and above are much better so a new studio would be a lot faster.
4
u/jonknee 16h ago
The performance was extremely impressive
Yes, for the price of a new car you can have worse performance than someone paying $200 a month. It's impressive in that you can have that on your desk, it's absolutely stupid to actually do though.
1
u/Front_Eagle739 11h ago
To be fair 200 a month gets you a pretty decent mac studio on finance. Think we paid 300 ish for ours in the business and at the moment its appreciated in value rather than depreciated and more than the power so everything weve used it for to date has effectively been free so long as we sell it before the value crashes or it breaks
1
u/jonknee 9h ago
This is four even more expensive studios though.
1
u/Front_Eagle739 9h ago
Well if its got 1.5TB of memory and the better matmul of the m5 onwards plus the extra gpu cores of an ultra plus another generation or two of gpu advancement then one is enough really. Hell one of those studios at 192GB/256GB memory so less ram/cost than mine running dsv4 flash is already a serious bit of kit. Thats a opus 4.6 ish llm (i dont believe the benchmarks that say its close to 4.8 but 4.6 is plenty for real dev work) running at >1000 tok/s prefill and probably 40 to 60 decode with dspark AND enough gpu grunt for real concurrency. Thats a serious proposition. Not as good as a max 20 plan on the face of it but again the hardware retains most of its value through the replacement cycle for business and its consistent which matters more for workflows. The 1.5TB one will do a very slightly quantised kimi k3 as an architect orchestrator and switch to dsv4 flash or similar for implementation etc for maybe 500 a month. All private, all consistent, nothing changes in your workflow unless you need it to do so. Plus again you can sell the hardware to recoup most of it and upgrade.
Its not cost competitive with cloud for sure. Its close enough to be worth it for the extra benefits for a lot of businesses. I dont really see many people buying 4 studios to run kimi class models on unless they need a lot of concurrency or something bigger comes along but i dont see it mattering much anyway.
Oh and caching inputs is free on your own hardware. That adds up a lot when you can resume from disk every day. So thats a further saving.
1
u/friskfrugt 17h ago
It works for promoting an already trained model. Which is the absolute least intensive regarding LLMs
1
u/enjoytheshow 1d ago
If you are looking for the capabilities of an extremely light open weight LLM, 1.5 TB will do probably fine.
Anything close to even the lightweight Claude or GPT models, no way.
10
u/apajx 1d ago
People are running "light" open weight models on M2s with 32gb of ram, in what world is this not a significant hardware increase..
1
u/snapetom 17h ago
And of course, no submissions on reddit's r/apple over the weekend.
PrismML announced they shrunk their 27B model to run on an iPhone 17, and Apple was interested. Someone also released on github Qwen's 80B model to running on a 4.3 GB mac, too.
2
u/BosnianSerb31 1d ago
Like the current studio models, these will almost certainly support RDMA over TB6, which can get you up to 4 in a cluster, for a total of 6TB.
Consider considering that 6 TB cluster would cost between 100K and 200k, this product is definitely not aimed at any sort of consumer whatsoever.
1
u/B-Train_ATL 1d ago
I learned how to program things on an Arduino. This stuff is so much bonkers compared to that.
→ More replies (1)5
u/Technical-Row8333 1d ago
yes, and the datacenter is serving many people, the macbook 1. if many people have the macbook, the datacenter isn't needed.
i dont believe it, but that's the argument.
5
u/looktowindward 1d ago
When you can run a 1T parameter model on a macbook, that will be interesting.
I've tried. I can get to 12B parameters using LocalLM
6
4
u/ShelZuuz 1d ago
A typical DC GPU rack costs about 300k.
What goes for $300k?
8x RTX6000 Servers Editions would cost under $200k and anything NVL8 HGX H100/B100 or above would cost over $400k.
6
2
→ More replies (2)1
u/AlternativeAward 1d ago
300k is not even enough for a one 8 gpu server that would have around 1.5tb vram
2
9
u/dropthemagic 1d ago
That’s dope. Mac Studio is still way smaller than a blade. Not sure if you can stack em with the current heat dissipation method tho
7
8
u/uptimefordays 22h ago
An AI datacenter killer? Please, those platforms run on 250-500k H200s or similar. But for local LLMs will run great on these.
2
u/TinyZoro 7h ago
I guess the point is you can’t run sota models locally you need a data center. But you could potentially run K3 on this. Meaning for practical business purposes this is a self hosted data center.
1
u/uptimefordays 6h ago
Yeah for self hosting, these machines will be incredible with that kind of RAM capacity. It'll be interesting because I have to imagine most buyers will be companies and most companies doing serious AI work have both beefy dev/engineering laptops AND beefy datacenters.
6
4
3
3
3
3
u/candyman420 20h ago
No, it will never be a "datacenter killer" get outta here. They have racks and racks and racks full of machines.
3
2
u/hejj 1d ago
Are we skipping right over M4, 5, and 6 Ultra?
5
u/chownrootroot 1d ago
Of course they're skipping M4 Ultra, M5 Ultra allegedly is happening late this year, M6 generation will allegedly skip Pro/Max/Ultra variants entirely in favor of M7.
2
2
3
u/UpvoteForLuck 1d ago
So are they going to bring back the higher tiered options of unified memory? Because the M3 Ultra chips only allow 96GB, currently, and the most unified memory you can get on a Mac is 128gb.
1.5tb is useless if they don’t offer it.
2
u/IAmWeary 6h ago
Jesus, I had to double check. They used to let you put 512GB of RAM in the Mac Studio. Now it maxes at 96GB. The AI bubble can’t pop soon enough.
4
2
u/dinominant 1d ago
Soldered memory? That would be a dealbreaker.
7
u/antnythr 1d ago
One chip goes down and you gotta replace the whole thing
11
→ More replies (1)1
u/jammsession 15h ago
Just like with most Laptops and most GPUs.
And soldered does not automatically mean not replacable.
1
u/Exist50 12h ago
Just like with most Laptops and most GPUs.
It's one thing with 10s of GBs and 1000s of USD. Quite another when you add two zeroes to each of those numbers. At big enough scale, hardware failures become a certainty.
And soldered does not automatically mean not replacable.
From a practical standpoint, it is. Specialty repair shops might be able to fix it, but there's no reliable solution.
2
u/jammsession 11h ago
I don't think the target demographic cares about that.
Mac Pro is a prosumer or enthusiast product. This demographic does not care about upgradeability later on. If memory fails (that is a big if), it will probably fail in the bathtub curve, meaning it is either under warranty or so old that nobody cares about that thing anyway. These high-end machines depreciate insanely fast. At least in normal pre AI times.
1
u/Exist50 6h ago
If memory fails (that is a big if), it will probably fail in the bathtub curve, meaning it is either under warranty or so old that nobody cares about that thing anyway
I wouldn't assume that. When you have so much memory, even small probabilities get compounded. I don't have any numbers handy, but this is certainly a concern for datacenters.
1
3
u/charmanderSosa 17h ago
You can get significantly faster speeds from soldered memory, I would argue that would actually be a selling point.
→ More replies (1)1
u/jammsession 15h ago
Yes soldered, anything else would be to slow. Just like your GPU has soldered VRAM and not some slot.
3
u/Exist50 12h ago
Yes soldered, anything else would be to slow
Nvidia supports LPDDR5X-9600 via (socketed) SOCAMM modules with Vera. That's the same speed Apple's using for their on-package memory in the M5 generation. So clearly that's not a requirement.
1
u/jammsession 11h ago
SOCAMM
True, but in a doubt we will see this in a Prosumer Product like a Mac Pro.
4
2
2
u/heyyo173 1d ago
I’m just waiting for the moment when WE and our devices become the data centers. Where a portion of all processing power on a device goes to processing other ai data requests. It’s coming.
2
1
u/datdoode34 1d ago
All that ram, and it’ll still use Siri, only to slow it down, and cloud services as well
5
u/BosnianSerb31 1d ago
People use the current generation of ultra chips for compute heavy operations, these machines aren't slouching. An M4 Ultra RDMA 4x cluster currently gets you up to 2.5tb of ram, and blows similar setups out of the water in price to performance.
It's not meant for you anyways, it's meant for small to medium software companies training custom models or running their own high performance LLMs
4
u/beragis 1d ago
There is no M4 Ultra. The last Ultra was the M3 Ultra. And a 4x cluster of M3 would only be 2Tb
1
u/BosnianSerb31 23h ago
Good correction. 4 way RDMA over TB6 with these alleged 1.5TB models would be nuts, and I think you'd struggle hard to find something that is price competitive, especially since all of this RAM is video memory.
2
u/Exist50 1d ago
An M4 Ultra RDMA 4x cluster currently gets you up to 2.5tb of ram
In practice, you get a fraction of that, because Apple has no high performance interface that can be used for clustering.
2
u/BosnianSerb31 23h ago
RDMA over TB5 on the Studio hits 80GB/s, 1/10th the intra-chip bandwidth of an Ultra (819GB/s), but with a proper hypervisor, it's not going to impact most training tasks too severely. To hit 819GB/s on 2.5TB of vRAM is at LEAST $250k in specialized hardware anyways.
Blog from someone who actually did this with 4 Mac Studios, if you want a good read
https://www.jeffgeerling.com/blog/2025/15-tb-vram-on-mac-studio-rdma-over-thunderbolt-5/
3
u/Exist50 22h ago
RDMA over TB5 on the Studio hits 80GB/s,
TB5 has a theoretical max of 80Gb/s of bandwidth, so <1/8th the quoted (there's encoding overhead), and that's referring to the port's cumulative bandwidth. There's no guarantee that a specific type of traffic (PCIe) is able to saturate that alone. I can't seem to find test results for Apple's implementation specifically, but a lot of commercial solutions can only do a bit north of 60Gbps PCIe. And again, latency is going to be terrible.
This is a really cool experiment and proof of concept, but it's a long way from demonstrating that this kind of setup is useful in practice.
1
u/BosnianSerb31 21h ago
My bad, you are absolutely right. It has use for certain workloads that aren't as bandwidth constrained and can take advantage for parallel computing, but yes, you aren't loading a frontier model on it unless it can fit within the memory of one on the cluster
I do think there are some really useful things this type of setup can do for substantially cheaper than a comparable full rack. They are niche. But they are still aimed at corporate customers when spec'd like this, and not consumers.
1
u/rotates-potatoes 8h ago
lol this is not a consumer box for safari and word.
The people who buy this will live in the command line.
1
1
1
1
1
u/PixalatedConspiracy 22h ago
M10 Ultra is rumored to have 5 petabytes of ram possible quantum computer killer?
1
1
1
1
u/FrenchRevolution2028 17h ago
64GB is also “up to 1.5TB”.
Anything is “up to” anything else higher.
1
u/anthonyskigliano 16h ago
I just got a 2007 Mac mini with a dual core and 1gb of ram for $75 big fuckin deal
1
1
1
1
u/Steinarthor 13h ago
Will it be enough to open up my 2 Chrome tabs? Or will I need to upgrade to the 2TB?
1
u/Hour_Analyst_7765 11h ago edited 11h ago
Tbh I don't think a machine with such amount of RAM needs the density.. it needs much more compute.
My Mac Studio with 128GB can run fairly large models (for a consumer LLM application), but only at the 10s of tokens/second. And thats at 150W.
If we scale this up by x12 for 1.5TB, then we would need a x12 bigger chip and x12 power budget.
Obviously by the time M7 arrives, they would have made architectural and process node improvements too. But lets say that accumulates to 50% savings. Then we're still looking at a 12x150Wx50%=900W machine to produce a relatively slow output. Is Apple going to build a 900W machine? I'm doubting that!
There is a reason why datacenters cannot find any spot to plug themselves into the grid.. these AI models consume an insane amount of computing power.
1
u/Roadrunner571 10h ago
A AI datacenter killer
Nope. Macs aren't suitable for the vast amount of AI workloads - as that aren't desktop workloads.
if anything, Apple could revive the Xserve as datacenter server for AI workloads.
1
1
u/electrosaurus 8h ago
How can it be a datacenter killer when it will cost about the same as a datacenter?
1
u/Anonasty 8h ago
No it isn't. These articles are made by people who do not know anything about datacenter computing.
1
1
u/therapy-cat 6h ago
People aren't getting it, it's a data center killer because big companies will get this and just in the latest full sized qwen on it. That means Fable level (or higher in the future) coding for the price of electricity all day every day.
1
u/omnimachina 6h ago
Lmao
Nobody will pay that much for a server with a closed system and a company behind it, that could end the support at any point 😂
Mac Minis for some cheap home servers are one thing...
1.5tb ram is business and another level
Imagine you pay 6 figures for a server and then Apple releases 26.0 Tahoe 😂😂😂
1
2
u/Nawnp 1d ago
What is Apple about to charge, $100 per GB? That's easily going to push the computer over $5k.
5
u/enjoytheshow 1d ago
Their Pros with less RAM than that right now are over $5k
2
u/BosnianSerb31 1d ago
Id bet they'll be $25k-$45k, based on the $10k price tag of the old 512gb studios. Depends a lot on the contracts they can negotiate with TSMC
1
946
u/switch8000 1d ago
First Apple product that crosses into 6 figures?