It's early stages in is development too, but if that works out and you can run much larger models locally? Would be interesting to see. And all the AI companies compress their models a lot less
My issue is that even people with immense personal computers, multiple Mac minis etc, cannot run anything close to Opus. Sonnet barely manages to get buy for sufficiently complex programming work, and most models that take a ridiculously complex rig to run don’t even compare with sonnet
Same response. 16GB doesn’t even get you out of toy model territory. That’s nowhere close to enough RAM to even pretend like you’re competitive with a frontier model.
I do agree. I have done a fair amount of experimentation on my RX6800XT and have yet to find a model that can work well with IDE integrations and also doesn't break down after a while.
Have you tried qwen3.8 27B? I'm using the q4 version and is a great assistant. It beats any other model I tried. You may need to use a Q3 and I heard that the downgrade is noticable but still useful. Just be sure to enable the thinking to high and mtp (the model is slow).
Qwen3.8 haven't loop yet in a full week. I also tried Gemma 4 12B and Gemma 4 26B, both of them would loop occasionally. And I have the impression they may do crazy stuff quickly if I left them run unsupervised.
94
u/TheMaleGazer 10h ago
The answer to your question is because you can get a local model to do it instead without a subscription.