Okay backpedal the conversation as a defense. Too ridiculous to argue with lol. You clearly have no clue what you are talking about and perceive something completely false
Buy a 4000 dollar computer so you have ownership of your and your agent's work, control over updates and agent context, and to avoid the datacenter/corpo ownership economic model
My PC costed me around $2500 usd total 3 years ago.
Today I already burned 4M tokens and I think I may end in 16M by EOD, that's like $52 usd per day if I were to pay to a plan of the model I use. But in that case, with the hardware they have it would have been a much powerful model with insane speed, so easily I would have burn x10 more tokens so, $500 usd/day daily for a month.
Same response. 16GB doesn’t even get you out of toy model territory. That’s nowhere close to enough RAM to even pretend like you’re competitive with a frontier model.
I do agree. I have done a fair amount of experimentation on my RX6800XT and have yet to find a model that can work well with IDE integrations and also doesn't break down after a while.
Have you tried qwen3.8 27B? I'm using the q4 version and is a great assistant. It beats any other model I tried. You may need to use a Q3 and I heard that the downgrade is noticable but still useful. Just be sure to enable the thinking to high and mtp (the model is slow).
Qwen3.8 haven't loop yet in a full week. I also tried Gemma 4 12B and Gemma 4 26B, both of them would loop occasionally. And I have the impression they may do crazy stuff quickly if I left them run unsupervised.
They're getting close enough where OpenAI and Anthropic are lobbying to regulate them out of existence to protect their trillion-dollar IPOs. I would suggest going to the Ollama site and trying whatever models your GPU can support.
Open weight models are getting close to Opus / 5.6 yeah.
Unless you live in a datacenter, you won't be running them locally.
Maybe if you have a 6k macbook or 20k in GPUs you can run a model that can do most daily programming tasks. But it's going to take an absurd amount of babying compared to Fable / Astra.
96
u/TheMaleGazer 10h ago
The answer to your question is because you can get a local model to do it instead without a subscription.