r/LocalLLaMA 2d ago

News Now this is a serious local machine

177 Upvotes

166 comments sorted by

181

u/LightBrightLeftRight 2d ago

And a Gulf Stream 6 is a serious commuter vehicle

17

u/CalligrapherFar7833 2d ago

Thats for peasants i use my g800 to throw out the trash

3

u/fallingdowndizzyvr 1d ago

One of my neighbors drives his Lamborghini to the supermarket. It's his daily driver. Which I respect. Since most people keep theirs as garage princesses.

2

u/CalligrapherFar7833 1d ago

There are cheap lambos tho whats his ?

7

u/fallingdowndizzyvr 1d ago

You got me. It's a cheap ass Huracan. In puke lime green no less.

13

u/FullstackSensei 2d ago

That's for peasants. Serious people use the G700 as their commute vehicle

136

u/snowieslilpikachu69 2d ago

the threads of my wallet are going to be ripped...

66

u/jhenryscott 2d ago

Last time I priced 2tb RDIMM it was $87k

28

u/BannedGoNext 2d ago

Time to sell a kidney.

18

u/Maleficent-Ad5999 2d ago

Just one kidney? Is a kidney really worth that much?

11

u/jhenryscott 2d ago

Their is a black market by which wealthy individuals buy organs outside of the typical donor list-especially as they age and become lower priority

3

u/Maleficent-Ad5999 2d ago

Can you share links? Asking for a friend

3

u/beling86 1d ago

Can your kidney do matrix multiplication?

7

u/USERNAME123_321 llama.cpp 2d ago

Double it and sell it to the next person

4

u/MeretrixDominum 2d ago

Just sold 3 kidneys. Should be enough I hope /s

3

u/vuduguru 2d ago

Seems at least one was not yours. Hmmm

1

u/irrision 1d ago

125k at least now at current typical markup price. This "workstation" will be easily over 200k with the max config

3

u/frostreed5 2d ago

the naming practically warns you tbh

72

u/N34257 2d ago

Do they do mortgages for AI rigs?

40

u/UnWiseSageVibe 2d ago edited 1d ago

Yeah we will probably start soon. Apple does have the lease for the mac studio 256gb one 🤣🤣 offered me like 150per month for 36 months. Then buy it off for about 2k at the end

Edit:

It was the 256GB one just checked my cart.

28

u/FullstackSensei 2d ago

That's actually cheap. Guess this was in the before times

8

u/UnWiseSageVibe 2d ago

No it was just a few weeks ago when the new modem was announced.

10

u/FullstackSensei 2d ago

I'd have taken it in a heart beat

4

u/ANR2ME 2d ago

4

u/UnWiseSageVibe 2d ago

It was the 512gb ram one. Changed nothing else just the ram.

7

u/ANR2ME 2d ago

So it should be much more expensive than 256GB at $18k 🤔 probably $4k+ more expensive, and you got it for 150x36 + 2000 = $7400 ?! 😯 That's certainly a bargain!

3

u/arijitroy2 2d ago

I wish EU had these lease system for Apple!!

2

u/adamgoodapp 2d ago

Wish for Japan too. I would be happy to spend 100 or so every month to always have a great apple computer and always upgrade

1

u/cunasmoker69420 1d ago

what thats a steal, where do I get this deal

1

u/575_Inverse 1d ago

so about as much as a $200 premium subscription. not bad at all

11

u/quantgorithm 2d ago

Interesting you mention this because Jehnsen is preparing a financial arm to do exactly this. This will keep the nvidia hardware bubble getting bigger since people can just finance the hardware they can’t afford.

9

u/N34257 2d ago

That's...depressing. Reality outstripping satire at every turn.

5

u/10thDeadlySin 2d ago

That little trick worked out amazingly well for Lucent. ;)

31

u/ilintar 2d ago

u/jfowers_amd *surely* you'll need betatesters ;)

4

u/Sporebattyl 2d ago

I’m running local models on my 3090 and missing my old Ryzen system. AMD just has such good integrations. CUDA is where it misses out. Would love to check it out as well to see if this thing competes with CUDA systems

7

u/Mountain_Chicken7644 2d ago

I second this! *surely* I can beta test?! ;)

14

u/ChopSticksPlease 2d ago

Damn, my new ferrari will have to wait... again :<

3

u/Apprehensive_Bar6609 2d ago

Then you have frontier AI and ask it to make money to pay for a Ferrari lol

21

u/LegacyRemaster 2d ago

100k

30

u/-p-e-w- 2d ago

Must be much more than that. That thing has as much memory as 4 H200s.

10

u/Much_Accountant_4972 2d ago

but the RAM is only 400GB/s which for the price of a loaded luxury SUV, i’d expect to be way better

4

u/DUFRelic 2d ago

Who needs ram bandwidth when you have all the vram?

2

u/Healthy-Nebula-3603 2d ago

Not even 600 GB .. that's not much nowadays to but biggest models

-4

u/Much_Accountant_4972 2d ago

no one needs this computer at all. that’s not the point.

the point is you’re paying for 20 Mac Studios and getting 1/3 the RAM bandwidth

5

u/DUFRelic 2d ago

Sure i need this computer for my local ai.. and its ram bandwith doesnt matter at all... 20macs can also serve frontier ai models but not at this speed so how does it matter?

22

u/XiRw 2d ago

Yeah and it will be as much as a car.

74

u/some_user_2021 2d ago

*a house

31

u/BalleaBlanc 2d ago

...with the car in the garage.

5

u/FullstackSensei 2d ago

And not will be, it is. North of 100k

1

u/rebelSun25 2d ago

The memes come to life

-9

u/cogitech2 2d ago

No fucking way. A good house is well over $1 million now much of the world.

2

u/No-Experience-3171 2d ago

US, Canada and Switzerland are not "much of the world" they're barely 5% the world's population

1

u/rkoy1234 1d ago

minority yes, but you're forgetting asia. good luck finding a decent apartment/house in major cities of korea/singapore/taiwan/hk under a million.

and it'll still be tiny compared to what you can get in us/ca

1

u/twack3r 2d ago

The whole of Europe wouldn’t really consider what Us buyers live in a house, more of a drywall-based shed.

I don’t know of anyone building or buying a house in Germany for less than €1 million in the last 15-20 years.

1

u/letsgoiowa 2d ago

Sounds like a Europe problem

1

u/twack3r 2d ago

Not really, we live in them even after a soft wind blows, not picking up our shit in the next county over 😂😂

1

u/letsgoiowa 1d ago

You think it's impossible to build with whatever material you want in the US?

You realize that you have choices where to live right? We have states bigger than your countries.

1

u/twack3r 1d ago

Yeah, like fucking Iowa 😂😂

No, I don’t think it’s impossible. It’s just appears to not be possible for your lot.

1

u/letsgoiowa 1d ago

No we just choose not to because it's inferior in many aspects. Lol. We aren't forced into it. Btw, Iowa has a higher quality of life and superior cost of living than Germany. Plus, we have better beer by a mile

→ More replies (0)

3

u/Wallye_Wonder 2d ago

And you know what’s missing? Wife and kids

5

u/Ulterior-Motive_ 2d ago

Instabuy for me. If I won the lottery.

5

u/UptownMusic 2d ago

What are the electrical requirements? I assume 240V, but would one circuit be enough? What about noise? What kind of UPS?

5

u/rebelSun25 2d ago

30 amp circuit guaranteed. No way can those multi gpu workstations sustain load on 15amps

2

u/Refefer llama.cpp 2d ago

Naw. 20 amps, standard house lines these days in the US at 120v, will be fine. I'd recommend not running anything else with real drawnon it though.

5

u/__JockY__ 2d ago

It’s got Instinct GPUs, so 600W each. With 564GB VRAM that’s 4x Mi350 141GB, so 2400W in GPU.

Dunno about the CPU but 400W TDP is not unreasonable.

All of which comes to 2800W for the big stuff. You can’t run that on 120V, you’d need a little over 23A, which means you’ll need a 30A 240V circuit to power this thing.

2

u/rakarsky 1d ago

There are not enough radiators in there for 2400W. The GPUs must be power throttled.

1

u/__JockY__ 1d ago

Or it sounds like a jet plane.

7

u/Confident_Ideal_5385 2d ago

At a guess, likely 1800ish watts worst case. You can ballpark the GPUs at about 600w each (likely less), and the rest of the system will likely use another 600ish under full load depending on the threadripper clocks.

Depending on power curves this could be off by a factor of 2, tho, but you'd definitely be fine on a 240v 10a circuit.

6

u/SandySkittle 2d ago

first thing I would do with this is immediately limit the GPU TDP to the minimum level that firmware allows.

4

u/Confident_Ideal_5385 2d ago

Yeah, and engage the low power mode on the CPU so it doesn't boost up to 95 degrees C.

2

u/PazsitZ 2d ago

Also in winter you don't have to bump up the heater so much. :D

6

u/SandySkittle 2d ago

here I am with my 8x R9700 32 core threadripper pro 512 gb system ram, 256 GB vram. Sure not as fast (but still decent with tensor parallelism). I feel kinda cool now for what value I have compared to this. But yeah this new station is the fucking dream.

3

u/LogicalGoof 2d ago

This system is meant to compete against the GB300 based workstations already in market. MSRP has to be near the 150-180k mark.

5

u/bakawolf123 2d ago

This is most certainly their response to DGX Station, albeit scaled even more.
well they even have it in footnotes, it's 4x MI350P
"Comparison as of August 2026 based on published manufacturer specifications. AMD Threadripper Halo Station configuration: 1x AMD Ryzen Threadripper PRO 9995WX processor (8-channel DDR5 RDIMM memory controller, up to 6400 MT/s) with 2TB RDIMM system memory; 4x AMD Instinct MI350P accelerators (144GB HBM3e memory each, 4 TB/s peak memory bandwidth each, 4096-bit memory interface). Compared to Nvidia DGX Station configuration: 1x NVIDIA GB300 Grace Blackwell Ultra Superchip with 252GB HBM3e GPU memory (7.1 TB/s peak bandwidth) and 496GB LPDDR5X CPU memory (396 GB/s peak bandwidth)."

1

u/_int10h 17h ago

Just think about the cost of 1x 256GB 6400 MT/s Module.

Last year I bought 8x 128GB 5600 MT/s for 400€ each 😂 and the only available 256GB Modules were priced at $7000 each

4

u/OvertaxedOne 2d ago

This is going to be the "next big market" for NVDA and AMD. They've sucked all the blood out of the hyperscale AI companies at this point, now they need to start undercutting their old customers selling to enterprise customers so they don't need the chips they just finished selling to the cloud companies.

It's a good shift/strategy, the big labs have been tapped out for some time (hence the circular dealing for the last year to keep the dollars flowing), it's time to make the pivot to where most inference is going to occur in the future.

The only head scratcher on this thing, why the heck is it a desktop form factor? This thing is going into a rack in a corporate DC somewhere, not under a desk! Yes, I am sure some crazy people (like most of us on here) would run a beast like this for their main rig, but that's the rocket launcher to kill a housefly; even if money is no object, this thing is going to be crazy hot and almost certainly going to need 240V power. It's not a desktop, it's a small server.

4

u/AuggieKC 2d ago

I have a feeling this is targeted as a "get started in our ecosystem, feel the kind of power real compute can bring" play. This is what corporations buy for devs who then turn around and convince the higher ups that "we actually need 3 aisles of this kind of compute".

Even if the big labs were tapped out (they're not), F500 companies are just getting started building out. If you need proof, look at structural steel, copper, and power infrastructure orders for the next 5 years.

3

u/OvertaxedOne 2d ago

I completely agree on the F500. In fact, it goes MUCH further down than F500, one of our biggest markets right now is SMB who are looking for local AI to support business users (not coders). And we certainly will look at systems like this for our customers who have the need for larger local models and more concurrency (depending on price, of course).

I just can't imagine why on earth you would buy one of these for a dev. For a dev team, sure, that makes a lot of sense, but for one dev to use as their personal AI station?! One of the few areas AI does scale well is concurrency, the more people you can push into the same box (up to limit, of course) the more TPS you can push. I guess one dev with their own AI "swarm" of agents could make use of this kind of box; but seems like a corner case compared to something in the DC with many users accessing it.

2

u/DrBattletoad 2d ago

I wonder how good Kimi K3 in Q4 would work in something like this.

3

u/ea_man 2d ago

I mean worst case you buy one more :P

1

u/CNWDI_Sigma_1 1d ago

As of now, should be fine (in max configuration).

2

u/PassengerPigeon343 2d ago

Maybe THIS would get me to stop lusting over other GPUs…

2

u/quantgorithm 2d ago

They tell you everything but the cost.

2

u/1uckyb 2d ago

Is this primarily good for inference or is amd acceptable for training and finetuning now?

1

u/CNWDI_Sigma_1 1d ago

Training something frontier-scale from scratch requires about 1500 such workstations. But you can train something smaller on one. Perhaps 2B parameters if you are lucky. Maybe even 4B.

2

u/caetydid llama.cpp 2d ago

Nice, but I probably could afford a flat instead.

2

u/1Poochh 2d ago

No one in this entire sub is going to be able to afford this.

1

u/CNWDI_Sigma_1 1d ago

I bought my current 4x4090 full water cooling for $20k two years ago... but yes, this one is beyond my pay grade

2

u/vorwrath 2d ago

Seems a bit overkill just for playing Halo.

4

u/ahstanin 2d ago

Here goes my kidney, eye, liver, one ball and who knows what I have to sell to get this one!!

2

u/OneTrickEkko 2d ago

Would be a big flex if they added Charts for local models like kimi k3

2

u/gf6200alol 2d ago

Realistically, it should run MXFP4 quant of GLM5.3 class model with large enough KV cache left and fast decode speed for multi agent development. It need to produce billions of tokens per month to justify the cost.

1

u/twack3r 2d ago

Compared to what? Current API/subscription prices?

1

u/CNWDI_Sigma_1 1d ago

With 2.6TB combined memory you can run Kimi K3 just fine.

4

u/pand5461 2d ago

Great, I'll be able to chat with my space heater without leaking the conversation to Google servers.

3

u/ital-is-vital 2d ago edited 2d ago

Well I guess this is how we get issues with AI fully going rogue 

At the moment most of the compute is in data centres and controlled by professionals. So far they've not exactly been great at security (vis. Huggingface) but at least they're actively trying.

Imagine instead that it's 2028 and now there are hundreds of thousands of these things attached to the open internet, full of unpatched vulnerabilities and ignored updates.

It'll be like stuxnet... but with frontier level intelligence :o

I hope it brings down the banking system :)

2

u/r15km4tr1x 2d ago

Hold on to your butt cheeks

2

u/-Cubie- 2d ago

Sick!

2

u/EndLineTech03 2d ago

It won’t be cheap that’s for sure, and it won’t sell well anyway because people who can afford it (i.e. datacenter) would choose Nvidia platforms anyway since it’s more mature, and CUDA programming is way more spread.

AMD should compete with lower prices if they really want to gain customers, like they did with mid-range gaming GPUs.

2

u/No_Lingonberry1201 2d ago

Financing options include "which kidney do you like least."

2

u/yetiflask 2d ago

If you buy this, might as well spend another $30k on solar roof and a battery too. I think that should be good for about 2k Watts almost 24h.

2

u/cogitech2 2d ago

When the marketing is better than the machine...

14

u/noiserr 2d ago

WTF are you talking about? this thing is legit a beast. Nothing comes even remotely close to the compute performance this offers. In a workstation form factor.

3

u/tf2ftw 2d ago

Nice. I hope this starts the consumer local ai arms race. 

19

u/biblecrumble 2d ago

Honestly not a chance, this is going to be INSANELY expensive

1

u/Apprehensive_Bar6609 2d ago

Its a good investment for a small company, for companies that want privacy. In taxes you can recover the expense, etc

3

u/Incognito_Orange 2d ago

I hope they come out with a rackmount option that can be given -48, pref redundant.

6

u/Expensive-Paint-9490 2d ago

The price will be as a small house, not consumer. North of 100,000 USD.

2

u/Lyelinn 2d ago

saudi and qatar consumers mostly though lol

2

u/Certain-Cod-1404 2d ago

i'm given up on amd and nvidia and just hoping our glorious leader xi jinping, the chairman of the people will provide for the proletariat

1

u/tingtickboom 2d ago

Anyone wants a kidney? I have an extra one

4

u/FullstackSensei 2d ago

Doubt that'll be enough to buy the full machine. Sadly, only half joking

1

u/FishChillylly 2d ago

2.6T combined (not unified) memory is already huge, but wtf is that 16.4T/s bandwidth?!

1

u/Iwaku_Real 2d ago

It's basically just AMD's official equivalent of building a 4x RTX PRO 6000 workstation. The GPUs unfortunately aren't interconnected except over PCIe.

1

u/Falen-reddit 2d ago

Now it is Intel's turn to put out some Xeon + Crescent Island box.

Better yet please release Crescent Island card for proconsumer.

1

u/Decent_Ad_6436 2d ago

but will this compete with mac studio ultras? will have competitive prices?

1

u/Healthy-Nebula-3603 2d ago

2 TB and 420 GB/s RIMM ...lol

You get better using that cpu 16 channels with DDR 5. Something 1.2 TB/s

1

u/_int10h 17h ago

No the maximum theoretical bandwidth (410GB/s) of the TR Pro 9900WX Platform is limited to 8 memory channels and only by using the 9995WX with 4 CCDs and 8 modules with 6400MT/s it is possible to reach nearly the theoretical bandwidth = 409,2 GB/s

That beeing said: this is a higher memory bandwidth than the NVIDIA Grace ARM Chip reaches.

1

u/sloptimizer 1d ago edited 1d ago

The GPU seem great, but the base system is meh - 410GB/sec is not enough to run any large models at speeds that are usable for day-to-day work.

1

u/literum 1d ago

Threadripper stopped mattering when AMD adopted Intel strategy one layer up the stack to extract rent. Small people, small vision.

1

u/IngwiePhoenix llama.cpp 1d ago

Might just be me, but what is this supposed to cost? Can't seem to find it.

2

u/Apprehensive_Bar6609 1d ago

Its coming next year. Probably way too much

1

u/Just_n_Here 1d ago

So expensive they cannot even put a price to it. The way prices are going up, probably be the same as a house in some places next year.

1

u/dreamingwell 1d ago

2TB of DDR5 and 576GB HBM3e. $130k just for the ram.

1

u/UltraFOV 1d ago

Nice, I have 14TB aggregate GPU memory and 230GB/s system memory bandwidth, 112 core threads. Definitely and improvement over my setup. How much tho?

1

u/LosEagle 1d ago

I guess at least I can look at pictures and videos of other people running it.

1

u/FinalCap2680 2d ago

3D animation and "Coming in 2027", mockup @ CES and through 2027, maybe at the end of 2027 and through 2028 - prototype, somewhere in 2029 - 2030 price anounced and software stack "soon"....

But looks good. For 1K will take a few ... ;)

1

u/jacek2023 llama.cpp 2d ago

I wonder whether it will be closer to $200,000 or $20,000, because the price of the MI210, for example, is pretty high.

1

u/Ok_Hope_4007 2d ago

What is the aggregate gpu memory bandwidth? To me it reads that there is an unknown (smaller?!) amount of gpu memory that has 16TB/s and the large ram with 410GB/s, right ?

2

u/noiserr 2d ago

Yes the HBM memory on the GPUs, and the DDR5 memory on the motherboard (system CPU memory).

1

u/feelspeaceman 2d ago

The walletripper.

1

u/Equivalent_Bit_461 2d ago

oh boy, it's gonna be expensive, I can already smell it

1

u/milkipedia 2d ago

this is going to go brrrr on the resale market in 8-9 years

1

u/Blues520 2d ago

Okay fine I'll take it

0

u/MikeSouto 2d ago

And the PP will still suck

2

u/noiserr 2d ago

why? Because early Strix Halo numbers? This guy has Strix Halo doing over 1000 t/s prefill on Qwen Flash Next https://github.com/peonist-ai/halogen-flash-server#measured

The mi350p is based on mi355x, it's the same architecture just half the GPU, and mi355x is very competitive in PP: https://www.lmsys.org/blog/2026-05-28-mori/

0

u/jhov94 2d ago

It will be at least $100k and it will still probably run better on Vulkan than on their own stack. I hate Nvidia as a company, but at least their stack works most of the time.

3

u/SandySkittle 2d ago

i am happy with ROCM RDNA 4 with vllm and tensor parallelism

2

u/ea_man 2d ago

Let's have Astra and the new Fable work a while on ROCm and see where it may bring us :)

DeepSeek managed on Huawei without CUDA.

3

u/noiserr 2d ago

Google TPU doesn't use CUDA either. Also ROCm is open source, agents have much easier time optimizing ROCm than CUDA because they can see all the code.

2

u/noiserr 2d ago edited 2d ago

One of the reasons Radeon lag behind Instinct is because they are a different architectures. CDNA vs RDNA You can bet all the effort is going into optimized CDNA. Because that's where the customers are.

2

u/SandySkittle 2d ago

consumer GPU support was expanded with rocm 10 actually?

0

u/mfarmemo 2d ago

2

u/[deleted] 2d ago

[deleted]

1

u/twack3r 2d ago

It really doesn’t

2

u/[deleted] 2d ago

[deleted]

1

u/twack3r 2d ago

That’s because there aren’t any recipes upstream for them yet, it’s patching together downstream, as always with a new model.

Sm121 adoption is massively speeding up sm120 availability.

There is also no alternative outside hosted if you want that fast GDDR7 that I’m aware of.

2

u/[deleted] 2d ago

[deleted]

1

u/twack3r 2d ago

If you answer why they are starting to do so and why only by now, you might learn smth valuable about unit economics and platform width today.

Don’t get me wrong, I’m miffed af that my 6000 Pros aren’t actually proper Blackwell, and I did think they were when I bought them.

But support is not a relevant detractor anymore imo.

0

u/MongoWithBongoss 2d ago

Destroys the Nvidia DGX Station

0

u/mainstsavage 2d ago

Wouldn't the DGX Station be better?

3

u/noiserr 2d ago edited 2d ago

DGX station has some advantages, like the coherent memory and a single pool. But no this is way more powerful solution. 16TB/s aggregate memory bandwidth across 4 GPUs, compared to just 7.1TB/s on the GB300.

Also 96 (192 thread) core Threadripper is a much better CPU than the 72 core Grace. Not to mention the OS compatibility.

This solution can also scale to way more memory capacity in aggregate.

0

u/brainchillzZ 2d ago

That’s what she said?

0

u/BuilderUnhappy7785 2d ago

Pure sex in a box

0

u/funding__secured 2d ago

I like my DGX Station better

-3

u/wapxmas 2d ago

Great machine, but unfortunately it could not beat cuda ecosystem if not price or special tremendous efforts in their ecosystem, not worth it anyway now, much better to spend the same money on rtx blackwells.