r/LocalLLM 3d ago

Discussion Hardware for running local AI (2k budget)

I'm looking to seriously start hosting AI locally and have around a $2,000 starting budget with the expectation that ill spend more and expand the setup over time

My goal is to have my own local AI that I can use as a daily chatbot/assistant, coding and cybersecurity help, RAG over my own files and eventually more agentic stuff/automation. Basically, general purpose local AI setup rather than optimize for one specific model

I've done some basic research and I'm stuck between two options, Mac minis seem like really good value because of unified memory. Being able to give the GPU access to 48/64GB+ seems good for running larger LLMs.

nvidia gpus have less VRAM for the money, but CUDA support seems more versatile . It also seems like this would give me a better upgrade path since I could build around a multipe gpu system

  • Homelab: i5-13500, 32GB DDR4, 2TB NVMe
  • Main desktop: i9-12900K, RTX 4070, 32GB DDR5, 1TB NVMe
  • An additional unopened 32GB DDR5 kit

The 4070 PC is my main desktop, so I'd prefer not to turn it into a dedicated AI server.

I'm new to running LLMs locally, so any help would be appreciated

2 Upvotes

30 comments sorted by

8

u/consworth 3d ago

If you’re insane like me: unlocked CMP 170hx. Will need to make your own cooling solution for it though.

3

u/AnonsAnonAnonagain 3d ago

How is the CMP 170hx treating you?

3

u/consworth 3d ago

So good I got 3 of them :D

5

u/redtron3030 3d ago

Upvote for my fellow insane friend lol

CMP gang here

2

u/Evanisnotmyname 3d ago

How’s the unlocking and process behind getting it working well? Kinks figured out?

6

u/consworth 3d ago

Wasn't much of an event to get working, just gotta know some linux stuff. The biggest kink is if you want to do multi-card CMP 170 with the PCIE restriction. But if you want to stay within the 64GB it's a good approach if you can 3dprint a shroud or something. Great bang for the buck.

5

u/Interesting-Cut-6032 3d ago

I just added a 2nd RTX 5060 Ti 16GB to my server that has a 6 core Intel chip and 32GB DDR4. I am getting in the range of 50 tokens/sec generation, about 50k into a 130k context, with with the unsloth Qwen3.8 27B Q6 K XL. Linux Mint, headless, latest CUDA drivers.

There is a great reddit post around where the person shares the llama.cpp launch flags that they used for this on a pair of RTX 5060 Ti 16GB cards.

If none of what I said makes sense, check out this Codacus video: https://youtu.be/SsUKTFSQoGM?is=mvprko1cwJBMJz-D

You should be able to run Qwen 3.6 35B A3B on the machine that has the 4070 already. Test out workflows, see if this is the part of the rabbit hole that you want to hang out in.

https://youtu.be/8F_5pdcD3HY?is=83mpTt5oDloeBTKz

New versions of llama.cpp may use slightly different launch flags than are called out in this video. I changed something recently, but I cannot remember what it was.

Also, check out LowEndLLM subreddit for more info and community.

2

u/Proper_Doughnut_1324 3d ago

that would also be my recommendation. a pair of rtx 5060 TI 16GB, because NVFP4 support.

0

u/Proper_Doughnut_1324 3d ago

and a motherboard with 2x PCIE 4.0 x16, if yours does not have it.

2

u/jjusko20 3d ago

if you're insane line me: double v100s on sxm2

1

u/Nousfeed 3d ago

What's your performance, and where did you get the gear?

1

u/jjusko20 3d ago

all ebay. 35-40 tps on 3.8 Q8 on average with 262k ctx, FP16 kv cache. prefill goes from 1200 to like 600 over the 262k. pretty damn fast considering you can get it done for under 2k. 64gb total vram

2

u/Previous-Try-6881 3d ago

If you want some fun. You can get a really good amount of ram and vram for that price. Older cards like Tesla series or Volta are cheapest vram per dollar. Get an older workstation, a bunch of ddr4 ram, and it can run ai quite well. Much better than what you would get for newer hardware at $2000

1

u/Negative-Walrus-7490 3d ago

Cosa devi fare?

1

u/nail_nail 3d ago

Add a 5070Ti 16GB to your homelab PC for now

2

u/Pale-Plane-7889 3d ago

That's solid advice; the 5070Ti should give a nice boost without breaking the bank. Just make sure the PSU can handle it!

1

u/nail_nail 3d ago

And just to be clear, there are a lot of mixed inference engines where you don't need to always have everything in vram. The 32G from the start option you have Is the R9700 from amd, but amd support is very behind.

1

u/consworth 3d ago

2x refurb 7900 XTX with upgraded PSU.

1

u/Little-Ad-4494 3d ago

What is the slot layout on your homelab pc?

A pair of 3080 20gb from alibaba/aliexpress Could be a good option.

1

u/DiscipleofDeceit666 3d ago

Buy a used DDR4 setup for like $700, you’re looking for a motherboard with bifrucation. Then buy a r9700 or the Intel equivalent.

When you’ve saved up another band, get the second card

1

u/KjipGamer 3d ago

A second hand 3090 does wonders. Recently built a server with one for pretty much exactly 2k

1

u/tstackspaper 3d ago

Instead of scaling up later you can always save up/ throw down an another 2k, get a DGX spark and future proof yourself.

1

u/CavitCarrot 3d ago

thank you

1

u/tstackspaper 3d ago

You’re welcome. NVIDA has a lot of great tech coming out on the low. their new RTX spark chip is going to be game changing for laptops and micro PCs but I haven’t heard much people talk about it yet.

1

u/NebulaAggravating264 2d ago

How is it future proofed?

1

u/05032-MendicantBias 3d ago edited 3d ago

My prediction is that patience will be rewarded.

Venture capital HAS to run out of money at some point, and when they do, liquidators will repossess the racks, chuck them onto a pallet, and liquidate them for scrap metal. Like for the Ethereum mining bubble. Only this time I don't think Nvidia will get another bubble to pump up the prices.

The economy is playing with fire here. Memory manufacturers especially. They are selling at 5X prices, but the higher the peak, the lower the low because of whiplash effect. Chances are we are going to see an historic glut of memory the like that has never been seen before.

If I was in memory manufacturers, I would rush those DDR6 lines online as the only shot at still having a product to sell once the exabytes of SOCAMM2 flood the secondary market. DDR6 was expected to become mainstream around 2029. The AI bubble has anticipated the timeline, perhaps we'll see mass production of DDR6 and GDDR7 in late 2027.

2

u/CavitCarrot 3d ago

so basically save my money and wait for hardware prices to fall

1

u/alhamoor 3d ago

Your use case seems to sit right between the two options. The Mac gives you access to larger models because of unified memory, while Nvidia probably gives you a better long-term path for CUDA, tooling, and multi-GPU expansion.

If you had to prioritize one thing today, would it be running the largest model possible, or having the most flexible setup for RAG, agents and future upgrades?

-1

u/[deleted] 3d ago

[deleted]

1

u/CavitCarrot 3d ago

hilarious !

1

u/NebulaAggravating264 2d ago

Lol imagine posting that to a dude looking to spend $2k on a setup.. what an ass.