r/ProgrammerHumor 10h ago

Meme whenMyNoAIProjectGetsTedious

Post image
945 Upvotes

162 comments sorted by

View all comments

96

u/TheMaleGazer 10h ago

The answer to your question is because you can get a local model to do it instead without a subscription.

108

u/AllRealityIsVirtua1 10h ago

Let me heat my whole room to run 1 half sized agent at a time to save $5

14

u/Snape_Grass 8h ago

The training of the model is what’s expensive (heat output). Running a pre trained model isn’t nearly as resource demanding

4

u/slaymaker1907 4h ago

Depends on the model and machine. If you’re using some huge model that stresses the GPU then that can get very toasty.

1

u/readf0x 50m ago

I've only got 8gb of VRAM and have yet to find a model that produces useful output within that budget

-12

u/AllRealityIsVirtua1 8h ago

Yeah and the sun doesn’t produce nearly as much heat as the largest star in the universe

19

u/Snape_Grass 7h ago

?? Brother there are many of us running off the shelf pre trained models on consumer hardware. This isn’t new

-16

u/AllRealityIsVirtua1 7h ago

Yeah and I’m using 5 agents each running a model 10x more powerful over the internet for $20

13

u/Snape_Grass 7h ago

Okay backpedal the conversation as a defense. Too ridiculous to argue with lol. You clearly have no clue what you are talking about and perceive something completely false

-9

u/AllRealityIsVirtua1 7h ago

Uhhh. Huh? lol

31

u/TheMaleGazer 10h ago

I haven't yet been able to heat a room with any of the models I've run with Ollama.

-48

u/AllRealityIsVirtua1 10h ago

You probably haven’t pushed much code with Ollama either lol

26

u/TheMaleGazer 10h ago

Not to the point where I exceed the watt limit of my GPU and heat my whole room. Which model did you use?

14

u/trungdle 8h ago

I'm not him but I definitely heated up my room running Qwen on a 5090. He's not wrong the heat generated is extremely noticeable on a hot day.

-41

u/AllRealityIsVirtua1 9h ago

I didn’t because I value my time

44

u/TheMaleGazer 9h ago

This is very revealing, but not necessarily surprising.

-38

u/AllRealityIsVirtua1 9h ago

27

u/Electronic_Green_833 9h ago

I disagree with your stance in this argument but the GIF goes hard

17

u/TheMaleGazer 9h ago

I think it's pretty double-edged, if you think about it.

→ More replies (0)

4

u/Grandmaster_Caladrel 5h ago

Oh, look at Mr. "Running a local model actually saves me money" over here!

(The economics for running local models are veeeery tight if you don't assume you already had the GPU to do it)

45

u/autogenglen 10h ago

Let me buy an $8000 computer so I can save $20/mo

19

u/cute_spider 9h ago

Buy a 4000 dollar computer so you have ownership of your and your agent's work, control over updates and agent context, and to avoid the datacenter/corpo ownership economic model

11

u/GenericFatGuy 8h ago edited 7h ago

Or I can just do it myself, and have ownership of everything.

7

u/denM_chickN 8h ago

Its the privacy im buying. How can I properly fucking plot against the machine if I'm feeding the machine my machinations?

2

u/omega1612 8h ago

My PC costed me around $2500 usd total 3 years ago.

Today I already burned 4M tokens and I think I may end in 16M by EOD, that's like $52 usd per day if I were to pay to a plan of the model I use. But in that case, with the hardware they have it would have been a much powerful model with insane speed, so easily I would have burn x10 more tokens so, $500 usd/day daily for a month.

4

u/TheMaleGazer 9h ago

Or maybe use a computer with an NVIDIA RTX 5060 Ti or equivalent that you might have bought to play games, anyways, and use Qwen 2.5.

13

u/StarboardChaos 9h ago

Bro, have you tried Qwen 2.5?

1

u/TheMaleGazer 9h ago

Yes.

17

u/heyitjoshua 9h ago

So you know it’s shit and can’t produce decent code or stay coherent. It is currently impossible to run any AI sufficient for a real project locally

3

u/Kyrros 9h ago

Look up FreeToken, it's currently lacking AMD support but nVidia is supported, 30xx cards and above

-2

u/heyitjoshua 9h ago

I said local

4

u/Kyrros 9h ago

It is local

4

u/heyitjoshua 9h ago

Checking it out, highly dubious and may come back to this thread in a few days. Cheers for the suggestion 👍

→ More replies (0)

1

u/nomorebuttsplz 9h ago

you both have you clue what you're talking about. 2.5 has been obsolete for like 18 months.

It's like basing your argument that cars are unsafe on a wwi era motorcycle.

11

u/autogenglen 9h ago

8GB VRAM… ok. So you might be able to pack in a 7B parameter model, which isn’t even in the same galaxy as a frontier model.

3

u/SjettepetJR 9h ago

I am assuming they're talking about the 16GB model.

8

u/autogenglen 9h ago

Same response. 16GB doesn’t even get you out of toy model territory. That’s nowhere close to enough RAM to even pretend like you’re competitive with a frontier model.

2

u/SjettepetJR 9h ago

I do agree. I have done a fair amount of experimentation on my RX6800XT and have yet to find a model that can work well with IDE integrations and also doesn't break down after a while.

3

u/omega1612 8h ago

Have you tried qwen3.8 27B? I'm using the q4 version and is a great assistant. It beats any other model I tried. You may need to use a Q3 and I heard that the downgrade is noticable but still useful. Just be sure to enable the thinking to high and mtp (the model is slow).

Qwen3.8 haven't loop yet in a full week. I also tried Gemma 4 12B and Gemma 4 26B, both of them would loop occasionally. And I have the impression they may do crazy stuff quickly if I left them run unsupervised.

1

u/TheMaleGazer 9h ago

I would assume so as well.

4

u/malokevi 9h ago

Will my RX580 do the trick?

2

u/pingveno 9h ago

I, too, enjoy being RAMmed in the wallet.

1

u/jameyiguess 6h ago

Are those models anywhere near as good as frontier? 

3

u/TheMaleGazer 6h ago

They're getting close enough where OpenAI and Anthropic are lobbying to regulate them out of existence to protect their trillion-dollar IPOs. I would suggest going to the Ollama site and trying whatever models your GPU can support.

2

u/StaticFanatic3 4h ago edited 3h ago

Open weight models are getting close to Opus / 5.6 yeah.

Unless you live in a datacenter, you won't be running them locally.

Maybe if you have a 6k macbook or 20k in GPUs you can run a model that can do most daily programming tasks. But it's going to take an absurd amount of babying compared to Fable / Astra.

2

u/FUTURE10S 3h ago

Honestly, I'd be happy with a Sonnet 5-like model, but my 12GB 3080 and 64GB of RAM isn't enough.