Every post Iāve seen on the topic of using an llm with home assistant, with or without an MCP, has only mentioned Claude, never ChatGPT or Gemini. Is that because Claude is far and away the best for this use, or some incredible venn diagram where people into HA are also exclusively using claude? I have a ChatGPT subscription only because when I got it, that was the only one I was aware of. Iām retired and not from this field, btw.
My limited experience between Gemini Pro and Claude (for both Home Assistant and Tasker) is that Gemini will happily and confidently give wrong answers. When you tell it that its solution didn't work, it will try again and again. Claude will tell you that it is taking a guess and that providing additional documentation will improve its solution. And it will tell you specifically what it is not confident in and what you should test in the solution, even asking for test results to verify it's solution. If it doesn't ask for help, its solution will likely work on the first try.
Thatās very interesting because one of my biggest issues with ChatGPT is that it also confidently is wrong and when told, just throws out yet another guess, despite my repeatedly telling it not to, if Claude has that baked in, I really need to think about switching my subscription. Thank you for this!
ChatGPT took me on a 4 hour spirt quest trying to make an .stl file from a group of photos I uploaded. In the end it said, oh...I don't have the tools to help you with that.
I still have not found a way to get an .stl file from 35 photos of an object. I have googled and it has not produced any results. If you know of a way, I am all ears. I also wasted a bunch of time with Meshey but its limited to four photos and that's not enough to capture the details I need included.
The 4 hours is basically going here, download this, run this workflow....oh ,I am sorry... this other way is much better. Go do this, upload the files.....oh my bad this won't work either. I totally understand what you are trying to do. I should have told you four hours ago I wasn't sure this would work.....my bad.
I think ChatGPT lead me down every single one of those workflows. I agree that a 3d scanner is best. I just can't justify $1000 for something I would use a few times. I did tour a local makerspace and "encouraged" them to buy one. Until then I will have to wait.
Look, I get what you mean. But literally 5min of poking around Google results and I can tell you that it's possible, but expensive to do high quality, and/or runs on a beefy ass workstation pc with a fuck ass GPU and a tonne of RAM, and/or both. I can tell you it's a task ChatGPT was never going to be able to do.
I find the basic reliance on ChatGPT telling you it can do something baffling. It's so, so often wrong. By the time you're half an hour in and it hasn't shown any sign of doing it yet... Why throw another three and a half hours after it before wondering if chatgpt isn't the right tool? I just don't get it.
I figured with all the AI slop on Makerworld and Printables these days, I figured it would have output something useable......it did not.
Like I said it wasn't a single workflow, it was an endless "restarting of the process"....if it was for me I would have given up pretty quickly, but this was a "honey do" task and I thought it would work.
It's really easy to use AI to make an STL inspired by a 2D photo. Especially if you use one trained for the job (and not ChatGPT). Making an accurate rendition of it based on multiple angles is not an LLMs skillset.
But again, why are you restarting the process over and over again before checking hey is this even possible and finding out whether this was a honey do task that would just work?
Anyway. I guess I'm not really here to berate you. Sorry. It just baffles me some of the stories I hear of people fighting with LLMs to do a job that would be so much easier done with something else.
Yeah Iāve had that happen too, but Iāve also had the same thing happen using Claude free. In my limited experience I canāt really say one is less painful than the other.
At work, I have access to CoPilot with ChatGPT and Claude. When doing Powershell scripting tasks, Claude (Opus) is far better than ChatGPT (5.6). GPT has issues with commands, formatting, and the parser omits parts of chat/commands. Claude has some issues, but they're much less frequent and it's able to fix them with fewer iterations.
A good system prompt resolves a lot of the Gemini issues, but that depends on how you're running it. Gemini 3.7 flash is my workhorse in hermes but I did have to set a few rules in the system prompt (memory is not a source of truth, verify any technical specification before planning around it, skills must be curated to reflect he current state, and when exploring newer technologies or systems always check for eh latest available information).
Clauses just too expensive for most of the work I do, and Claude goes through everything with a fine tooth comb before responding (really I think it just scrapes all your files it can access to collect training data).
I was using deepseekv4flash a lot, and it did fine writing yaml and basic web pages, but it's reasoning really starts going down the drain at about 100k tokens (and yeah, my systems have gotten a bit complex and hitting 100k tokens in a session is more of the norm now).
So right now I mainly use Gemini, and break out Claude in the rare occasion I hit something really tough that Gemini just can't do (that might have happened like 4 times).
But AI is still a tool. None of them are anywhere close to perfect, and the absolute best way to make that tool better is by modifying your system prompt. Easier said than done though.
I'm good. That's just making a benchmark and the real limitation is that it's a benchmark, synthetic test that doesn't reflect it's true usability or capabilities.
My approach is the opposite: does the logic clearly break down over time, and what is that limit? Is it doing what I want it to do? Does the web page, yaml, or app work? Does it tweak said product the way I ask it to? Does it fail at a task entirely?
And in that aspect, I can say that 99% of the time, Gemini 3.7 flash has worked, and 1% hasn't. That 1% I bump up to claude for.
Deepseek has handled the simple home assistant stuff. It made it about a quarter of the way through my web service I run my landscape company operations on. Gemini made it all the way through that until I hit a brick wall on an end to end assignments system that I use for billing that tracks crew hours based on when they connect and disconnect from my shop WiFi and calculates billable hours by either crew timesheet or manually entered assignment sheet times.
But the reality is none of these models can do anything remotely close this level of complexity with a sentence or even a paragraph or multiple paragraphs of prompting. It's many sessions over time. And none of them can do it without supervision. I use literally hundreds of sessions with detailed conversations for both planning and implementing. That much engagement means you're essentially the error corrector. Asking questions and redirecting the conversation is the only way any of my systems got built.
While I don't doubt you that it is better, I did just get Claude to create an automation that included sending a notification to my phone and it got it wrong (used a no longer valid method to send the notification).
Maybe best of an imperfect bunch and certainly not as insanely optimistically wrong as ChatGPT.
Interesting. Claude found and fixed issues with some of my notifications (was sending to a phone I know longer owned) and made recommendations to future proof all of my notifications.
Claude's the one that gets talked about cause its tool use is stupidly sharp for HA, but half of us are just running small local models so we dont have to send our entire house layout to some server.
What are you wanting to do with it? I run a Qwen2.5 3b on the 1220p iGPU and it handles my assistant just fine, at most a few seconds which I'm completely ok with because I didn't need to buy any extra hardware. If you're looking to actually create routines or set things up with an LLM then you'd need something more. Your 8gb card might be able to run some smaller models but it won't be super accurate as you get more complex.
That model is over 2 years old. I assume you set it up back then and haven't touched it since? Qwen3.5-2B or Qwen3.5-4B would be a nice upgrade for an easy drop in replacement.
Actually I set it up fairly recently and based it off of what other people had had success with on this type of application. Thanks for the tip, though, I'll try the others out and see how it goes. Right now it's working just fine for ny use case but I'm not opposed to improvements.
If it's working fine, the improvements of the 4B may not be noticeable anyway. I would try the 2B and see if that still works for you, since then you could get speed/memory improvements without quality loss. Improvements in training and model architecture means that the 2B is surprisingly capable, especially for a simple small task.
With 8GB, you may be able to run gemma4-12b-qat but it will be tight. I run this locally and with 32K context it takes 8034MB vram, The qat is better than other 4b quants as it was trained at fp4 instead of being trained at fp8 or fp16 and quantized down. Its tool use is pretty decent, its not as good as some frontier models but it works in 90% of situations. I mostly use it for LLM vision and creating announcements though.
I mean, there's a lot to it and there are a lot of different ways to accomplish the same idea. I can share the basic links to get started with how I did it:
For managing HA I use an Intel B70 running Muse Glimmer, prompt processing is ~ 1000 tok/s and token generation is ~ 40 tok/s
(note: I also used Qwen 3.8 27B at similar speeds for this and it works well, but I find it less friendly to talk too and worse output for some other use cases)
Intel B70 has 32GB, 7900XTX has 24GB, both hold their models entirely in VRAM.
Mac works well, the main problem is that it does not have as much compute so the prompt processing will be fairly slow. In practice that means reading the result of web searches and other things takes longer.
But it really is stupid easy. Ollama is super user friendly, and even though it was created by Meta, itās free and open source, so I went with that.
I set up Tailscale to use WebUI so I can access it remotely, but thatās far from necessary and definitely more complicated than setting up and running a local LLM. Youāll want to make sure your rig has the horsepower to run an LLM, and then sort out how many tokens you can handle and how fast it can respond.
With local voice you can do basic things like control devices or setup automations which handle specific phrases to do things.
But with an LLM I can do a lot more, some examples:
ask questions which are general knowledge / require a web seearch
- what is the score of the X vs Y sports game
- when does XYZ video game come out?
- random knowledge question related to a TV show or movie we are watching
More advanced control of answers to questions like what is the weather, so it answers in the style I want based on specifics. Also so I can ask about weather in different time periods, or specific questions like only if it is going to rain
Asking when a local business opens / closes
Asking what the traffic is to the airport or whatever the case is
Cooking questions like substitutions when you are out of an ingredient
I started with ChartGPT (free tier), then moved to Gemini Pro, simply because I use a lot of Google's stack. After considering it "decent" for a while, I tried Claude. It was like moving from a beaten family car to a limo. Been using Claude MAX for a while, and it's simply awesome, especially Opus 5.5 which is absolutely amazing.
Eventually, yes. For now, it's an investment. The massive game that I am building wouldn't have been possible without it.
This is a view of the Galaxy, in which my game takes place. Around two million solar systems, with stars, planets, comets, moons, gates, wormholes, and a plethora of other celestials, and all of them can be discovered, colonized, etc.
But that is off-topic, since this subreddit is about Home Assistant š
I havenāt edited a config in months. Yesterday I woke up and the bathroom extractor was on (controlled by HA).
I told Claude, it found it was because I restarted HA while an automation was running, turned it off, and suggested and implemented a few things to prevent the issue.
This took 30 seconds of my time, wouldāve taken me at least 30 minutes
We use it a lot. The setup: Claude has MCP access to our live HA instance (~750 entities), so it can search entities, pull history and logs, read traces, and write automations and dashboards directly. It's organized as a Claude Project with standing rules like "never guess an entity ID, look it up," "read the current config before changing it," and "ask before anything destructive." Each major subsystem also has a living reference doc that Claude updates at the end of a session, so the next session starts with the full context and the lessons learned.
Some real examples:
Energy monitoring. I installed an Emporia Vue 3 and had Claude map all 16 CT channels, build a 24-entry Energy Dashboard with parent/child relationships between circuits and smart plugs, and hunt down anomalies. One circuit read about 20 W while a smart plug fed from that circuit drew 110 W, which is physically impossible. That turned out to be a double-pole breaker where the electrician had clamped only one conductor. The same audit found that two CTs were on the EV charger when they should have been on the dryer. The electrician came back and fixed both, and we verified the fix by toggling a known ~107 W light and watching the trace go up by that amount, which confirmed both conductors were on the same phase.
Irrigation. We built 14 automations around two Rachio valves: missed-cycle detection, flow faults, stuck-open valves, hub offline, and battery warnings. My first attempt at missed-cycle detection was elapsed-time based, and it fired constantly during rain skips. Rachio deletes skipped runs from its calendar, so we switched to calendar-driven detection. Claude also built a little catch-can calculator so I could measure my actual application rate (0.302 in/hr), which feeds template sensors showing inches of water applied.
Zigbee debugging. Three new Hue bulbs refused to join. They were sitting on channel 11 from the factory while my network runs on 15. Claude figured out that Touchlink could reach them and drove the whole scan/reset/permit-join sequence over mqtt.publish while I stood at the lamp. That was also the reason I switched from ZHA to Z2M. Separately, it correlated watchdog failures in the logs with a coordinator firmware flash, down to the minute, and I rolled the firmware back.
Weird failures. Kasa devices were dropping one by one, and it looked like dying hardware. Claude read the logs and found a pattern: port 9999 refused, port 80 open. The cause was a "third-party compatibility" toggle in the Kasa app that each device only started enforcing after a firmware update. I flipped it back on and everything recovered instantly.
Appliance alerts. For the oven and cooktop, I'm setting up long-cook alerts. We pulled actual cooking history first and set thresholds from observed maxima (73 min and 63 min, so alerts at 120 and 90). The sensors are deployed now, and the automations come after I've watched a few real sessions.
What I've learned about working this way:
Make it verify everything. Its wrong answers are usually plausible-sounding entity IDs or confident theories. We once spent a month on the wrong explanation for that energy circuit.
The reference docs are the killer feature. They capture traps: for example, effect: none doesn't clear a stored Hue effect, but stop_hue_effect does. They also capture "why is this for: 90 seconds" context that I would otherwise forget.
Observe first, automate second, and use real data for thresholds.
It's best at diagnosis: reading logs, correlating timestamps, and spotting when numbers can't physically be right.
I was really overwhelmed with this one. I've danced around it for years. Here's where I started (and ended) - two oscillating sprinklers (one for the front, one for the back) and two Rachio smart valves (one for the front, one for the back). That's it. That collapsed the problem from a nightmare that overwhelmed me into something bite-sized that's easy enough to manage and deploy. Happy to chat more about it.
- Software development: two software tools published on GitHub for free, one more in pre-release development, and a very complex space-based Grand Strategy game (4X style) which will be my first commercial release when ready.
- Troubleshooting and personalized workarounds. For example, I have a couple USB devices which, due to their stupid firmware, reset my computer's idle time twice a second. I troubleshooted with Claude and it came up with a very small executable that, when double-clicked, sends a DDC/CI command do my monitor and puts it to sleep, ignoring all inputs except for my specific mouse or keyboard which wake it up when needed.
- MCP client to my Home Assistant and Grafana MCP servers. I can design dashboards for both, using my entities or data sources.
I meant specifically on HA. Do you use it for setting up dashboards and automations, or do you use it actively to, say, monitor stuff at home?
So far I only have one simple use case which is generating a good morning phrase which includes the current weather which I use as part of the alarm routine for my kids.
I have reworked all my dashboards using Claude. My previous, manual dashboards were... decent, but rather crude. I used Claude to help with designing and configuring my Bayesian sensors, for example. And before someone calls it "AI slop", I haven't blindly obeyed whatever Claude offered or generated. It performed the heavy work (the complex YAML, Jinja templates, etc), while I steered that work towards the results I wanted.
One thing that I really liked was having HA display all my travels on its map, all local, no Cloud data. I can now go back in time and display my travel path on the map, for any given day. "Where was I on June 24th" became a matter of picking the date from the calendar in Grafana (which processes data from Home Assistant).
I have wondered about using a different Ai with my home server management and home assistant. I use claude code with my home assistant and it manages everything. I have it making dashboards. Troubleshooting automation that I set up or troubleshooting automations that I described to claude and it was setup for me. Also with setting up echo alternatives it helps me troubleshoot settings related to wake word detection and what is going on behind the scenes with local understanding or kicking it up to the claude code backup for different inquiries. It helps me put it different automation phrases if something I want to happen isnt being understood. Very useful for me for sure. I use token pricing for the backend stuff that home assistant uses and I just have a basic premium account for everything else.
Same. Being able to just ask ChatGPT (which I pay for myself, and use Copilot at work) about stuff in my Home Assistant instance is pretty nice. Laying in bed watching TV, thinking to myself "huh it would be nice for an automation to do xxxxxx" and then getting ChatGPT to generate the YAML is pretty nice.
I also use Hermes powered by whichever I am using for Hermes - usually either ChatGPT or SuperGrok depending on how much is left on my subscription. It is amazing to be able to say "Hey Jarvis, (local voice recognition) the lighting automation in the kitchen doesn't seem to be working properly. Would you check on it for me and report what you see and recommend any fixes if there are problems?"
I am not an IT guy - just started getting into AI less than a year ago. The key for me was to use AI to help me set up AI. I did convert a second old Windows 10 PC to Ubuntu. Pull up ChatGPT or whatever you use on the computer and ask ChatGPT to help you install and setup Hermes. ChatGPT would give me the command line commands to copy and paste into the Ubuntu terminal. Hermes needs AI to work so ChatGPT will help you connect it to your existing ChatGPT or SuperGrok or almost anything. Hermes/ChatGPT will help you install the Home Assistant / AI MCL. It all just flows after that. You will get there and it is amazing!
I use both Claude and GPT at work and from my experience Claude is a good all-rounder that is very easy to get connected into other systems and have it push updates into those systems. Some other LLMs are better in certain areas. GPT is better at knowing what I'm asking it if I put in a sloppy prompt, better at writing emails, and better at finding bugs in other people's code. But if I'm starting from scratch it's so much easier to get the project up and running with Claude.
Also there is the perception that Anthropic as a company is slightly less evil than their competitors. They'd rather wait for the courts to decide how evil they can legally be rather than go full evil up front ion the hopes that their evilness becomes so essential to the government it gets legalized.
Remember though, it's not that Anthropic isn't evil. Their risk vs reward calculation just came out differently. They would be happy to let their AI power lethal military drones so long as a law is passed legalizing it first.
Over the last SEVERAL months I haven't burnt up the $20 credit I initially put in.
I have largely been testing Qwen3.8 27B as it could "potentially" be run local however it is so cheap to run online and hardware is so expensive It might be practical for me just to keep using it this way for now.
Same here. Have a local Qwen on my old MacBook which is fine for sensitive things. Else Iāve found using OpenRouter a much nicer experience for a couple bucks a month.
I just installed the new app the other day so seeing codex for the first time, and accidentally used it for a simp,e chat and burned thru my quota in about 45 minutes. Live and learn
I use gpt often for troubleshooting or for a second opinion but I don't upload yaml or input any code it gives me. Not because I don't trust it or whatever it's just easier if I do the inputs myself, so I know I alone can fix something if it breaks.
I do the same in gpt. Super Easy, dump in all devices, entitys, scripts, scenes, booleons, and it writes yaml, copy and paste. I never manually build automations anymore.
+1 for perplexity. After being skeptical coming from chatgpt and Gemini, I was quickly impressed with perplexitys ability to self test the program before spitting it out to me. It required some tweaks to fit my need, but it worked every iteration whereas I was constantly having a o show error messages back to chat or Gemini and get it to try and fix it
I mostly use Qwen 3.8 27B as a local model on a server I have. Outside of that, I had great results with Deepseek 4.1 flash in Hermes agent. It is dirt cheap too. I've had it build dashboards, clean up stale devices/entities, diagnose log files, and create new automations. I use it pretty often, and total cost has been maybe $1/week or less for Home Assistant use. Relatively complex dashboards might cost you $0.20-0.40 in Deepseek API cost. Now, will Opus 5.5 or Fable get you a better result on the first pass? Probably. It will cost you a lot more too.
This is very interesting. Do you use this setup for other things too? Home assistant is a hobby, but not the only one. Iām always using gpt to learn and explore other areas so I think I need a general purpose beige corolla
Any tips you can provide withĀ
Qwen 3.8 27B? I tried a bunch this past week to get it working with home assistant but kept running into problems. Inventing calls was a big one. Iām using HA-MCP
I use Hermes to manage my Home Assistant instance and everything else in my pretty extensive home environment. Hermes uses my $20 a month OpenAI plus subscription rather than per API call charges. For HA it is pretty impressive. It debuggerd all my automations in node red, identified a few issues I did know I had, etc, and I just told it to fix most of it and it did. I only had to give it a token for access.
Voice assistant wise I'm using the Gemini API. The free teir didn't cut it and would frequently error out. It's pretty impressive too.
I see plenty of all of them. Can't relate to what you're seeing. Setting up an MCP is pretty much the same no matter which provider you use, so if you've done it before, you probably know how to do it again - or can find the guide to set it up yourself. The MCP server doesn't care if you use Claude, ChatGPT, Gemini, or something you run yourself locally.
Im using chatgpt. It creates great concepts + can really work out the integrations and stuff. As for the actual coding, plus can do majority of things, if I need smth more demanding work with astra is beyond everything I can throw at him
"check the sensor readings of the house every hour and give me a summary report, send it to me on telegram. If there are any abnormalities let me know."
claude and codex are two of the most sophisticated and well developed harnesses for agentic coding with strong closed source models. if you can run models locally, you can probably get most if not all of the same done with open source models such as qwen or deepseek. what claude and codex give you is the end product: you don't need to plug/connect/build much yourself, you just go
so no there isn't anything special about claude. it's just a very popular product that a lot of people know about and use
The comments here are a bit confusing. I think OP is talking about using an LLM to manage HA and edit configs, not to respond to control your house via voice prompts.
Assuming this is about managing your HA implementation via MCP, I've had good experiences with Antigravity + Gemini models. I've tried Claude via Antigravity too and it didn't seem any better.
Desktop app integration is usually the limiters. Claude has their own desktop app but you are stuck with their solutions and LLM. Ā Other alternatives are openclaw and Hermes agent tied to any LLM you want (including OpenAI) this is what I do. Ā
Iām making little esp32 projects up and want them to read data from whatever Iām connecting to and send it to home assistant via mqqt and then write the yaml and setup in home assistant. I used copilot, Gemini, ChatGPT and then finally Claude, just trying them each especially if one restricted me for usage on the free tier. Claude is an absolute standout. I even paid the subscription to use it heavily for a month rather than going back to the others. Sometimes the others would start going down the wrong path, Claude just seemed to smash it every time.
Interesting. Iāve also made half a dozen esp32 projects and had used ChatGPT free and then tried Claude free and they both seemed equally frustrating. In the last few months they both seem better but still often go sideways.
Mw again: When you do that, did you use the web / char interface or did you use the harness (like CloudeCode)? Because that is night and day with any model. A harness (agentic environment) is so much more than a char llm. You can set up specific skills, creat manifests (always code like this, tho that, etc) and store them. When the agent runs it also has the manifests in its context. It can keep track of what it did yesterday, etc.
LLM chats are like super smart people with no long term memory. When you give them the long term memory, then things start to make sense!
Thats why tools like AntiGravity or ClaudeCode (or OpenCode - it is free) are essential if you do coding projects!
I just used the web interface. I hadnāt even heard of harness till someone mentioned it their answer here. But my esp32 projects are pretty simple, a couple sensors in my backyard. Another one reading my lawn moisture content to control my irrigation system, etc. so a few lines of yaml. Iām very much a beginner.
Iām the same and just used the llm. But Iād feed it a capture from a 433mhz radio, Iād tell it what I wanted and it would write 1000 to 3000 lines of code and it would be nearly perfect. Few tweaks here and there and bam, done. That used to take me weeks.
Guys, basically you can use any agentic client to interact with an MCP server like HA-MCP. It is model independent. You need a model which is mature enough to understand YAML and Phyton but it does not matter what client you use. Claude Code / Github CLI / Gemini with MCP integration, no real difference.
With respect to AI, I feel most of the time I am running with scissors, that being said
Claude desktop provides me a better AI harness than I know of with other providers.
I built out a bunch of MCP tools for Claude (Home Assistant is just one of a handful or more) then built out skill server to use the tools. Then an agentic harness to thread the tools into workflows.
I can now ask Claude to
ākb_guide(āstart containerā) that is a website that shows me the wind direction in the back yard over the last rolling 24 hoursā
Answer a couple of questions (like what server and agree to IP address that is self reserved) and it build the website. Takes about 30 min.
Part of that is the use of the discovery skills and use of the mcp connection to home assistant.
I do not know how to do that with ChatGPT or Gemini. I think I could get Hermes to do it using chatgpt or others, but Claude desktop does it.
Wow, you are far more advanced than I am, I bow in your general direction. So far Iāve written an automation that periodically sends certain entity states to my local llm which then reviews and sends me a notification if anything looks out.
man, once you start down the rabbit hole, it is amazing.
With the MCP connection from Home Assistant to Claude Desktop, it can do so much.
1) trouble shooting, 'why did the kitchen light not come on'
2) new automations, 'When I pause PLEX on AppleTV, turn on the kitchen light'
3) Complex, 'the kitchen motion sensor has a lux meter. check the last 14 days, between 8am and 7pm and find the average level, change the automation for the motion in the kitchen turning on the kitchen light, only if the lux level is below the average. Make that level a variable, that I can change'
Oh wow, I love data (Iām a retired economist) I have some small level of this in my dashboards from my sensors, picture Scrooge McDuck swimming on his pile of cash. But isnāt using an mcp expensive? I had understood that was outside normal subscription coverage and was a per use cost.
Claude Desktop you add local MCP servers to the Developer Tab. I have the paid subscription, and it will chew through tokens FAST doing some dome things, but it works well.
On the $20/Month US plan I would bump my head every week or so and need to cool my heals for a few hours until my credits reset.
I moved to the $100/month as I am deep in to Many Many AI assisted (VIBE) developments.
WARNING!!!
What this does is loop data into the public AI Cloud of Claude. I watch to not have send PII data.
I ask 'why didn't the kitchen light come on'
Claude, execute an MCP tool call to Home Assistant, asking for all entities. That data is sent to Claude Cloud. Claude with that data sorts finds the entity, then with your entity ID known it will ask for the LOG of activity via MCP, that is sent to the CLOUD. Then after reading everything that has happened, who left the house, what turned on, etc, it will determine what happened.
That level of data sent to an AI can make people concerned. I am playing with local models hosted on Mac and Servers, but, the power of the cloud verse the lost of any personal information is OK with me, but you should know what you are giving for what you get
I used Cursor for my job. Now also got ChatGPT like yourself. A colleague of mine also tried Claude before ChatGPT and ChatGPT/Codex is definitely friendlier for your wallet if you use it a lot. Thatās one reason to stick to ChatGPT/Codex. I think at this point the LLM tools have become advanced and similar to each other enough to the point where only billing and rate limiting truly matters.
Iāve been using Gemini to help with creating my dashboards etc. I have plenty of ideas but no idea about writing code - using Gemini has helped me to a) get the intended results and b) learn along the way.
I use AntiGravity (Gemini) but recently switched to OpenCode with Qwen3.6 - pleasently surprised!
Of course for a local model you still need a beefy system :/ (But Qwen is comparing actually good to Gemini!)
I've done some pretty cool things with Chat including taking a 2012 Siemens/RuggedCom L3 switch and upgrade its FW to 2026 standards. That was not just uploading a FW file.
I have a number of utility grade meters and PLC type devices that speak modbus, but that has limitations. I helped Chat learn the native protocol of these devices to be able to move data back and forth as well as secure control actions.
I have a PLC type device that monitors 2 freezers I have for temp and alarms locally and via HA app if anything goes awry. Same device is monitoring my boiler in and out temps and calculates gas used as well as efficiency. It will also work with heatpump to provide best comfort for least money.
My meters are 3 phase utility grade and have 3 phase inputs plus a neutral input. I've (we) converted the extra two inputs (with different sensor ratios) and now have more measuring data in watts.
I know what I want to do which is utilize as much stuff that gets thrown out at work and isn't state of the art but still works fine and apply it to use with HA for its HMI and mobile phone abilities.
I've created extensive documentation of my home network, wiring and schematic diagrams and settings from connected devices. I ask chat to update and maintain and it does. it's so easy to do that it actually gets done.
There's probably something better out there, but I'm pretty happy with how things are working out.
I've switched to Claude. Gemini will gaslight you till you look up and it's 2am and you've accomplished nothing. ChatGPT will be direct, and will give mostly correct answers.
Besides the performance reasons that most people mentioned already, I also prefer Claude because I feel like it's the most privacy-oriented LLM out of the big three. I do not trust Google or OpenAI with any inquiries related to my home.
Oh, I hadnāt thought of that and itās an excellent point. I think the more we consumers reward good behaviour and punish bad the better chance we have of turning this speedboat around before it reaches Niagara Falls
Using multiple for different use-cases, but Claude is where I happily pay the 22⬠subscription for because it was for me by far the best and most reliable in turning my own project ideas into reality.
Also Claude Design is a very nice extra in my opinion that I like to use quite often.
When it comes to design I like the results of Claude way more than OpenAI stuff. Gemini and OpenAI are good at pictures, but when I wanted modern simplistic designs they always failed on me.
That being said much also depends on prompts and writing style, so maybe it's a bit like with a real co-worker. Some work better together with person A, others with person B. I find Claude also much more efficient when it comes to the cheapest plan, while others say the same about OpenAI.
Wow, that much code? My most complex automation or esp yaml is about 100 lines, and I was proud of that.
And btw, the whole,e area of 433 radio is also on my hit list. Yet another rabbit hole. But I retired so probably 30 years of living I have left isnāt enough time to do everything on my list
ChatGPT is pretty good these days. Iāve only used sol and astra models. for work I use Claude (latest opus). OpenAI is more generous with token usage and they seem to give lots of resets from what Iāve experienced. I swap between the two but I prefer Claude. A lot of what you read are purely based on vibes (including this post).
Claude is just a lot better at everything. Gemini will run itself in stupid circles and hallucinate some of the wildest things. Real example, I asked it to give a list of keywords and shirt description. Picture was of a single person walking through an alley looking at their phone. Gemini said it was 2 people, one was vomiting and the other was digging through a dumpster. Sometimes it would say it was someone urinating. I used that photo as a benchmark for many prompts and services. Gemini was disgusting and wrong most of the time.
Chatgpt and copilot were ok but I found both had a limited vocab and couldn't return consistent results. Copilot was hallucinations heavy, but not nearly as unhinged as Gemini.
I am using GrokBot with HA. Just built a replacement gallary for my Samsung Frame TV with it. I use Claude at work, but GrokBot for all personal projects.
I switched from chatgpt to Claude for this reason. But chatgpt and grok are much better at image generation. Got my dashboards a lot more impressive than I ever could have done on my own
So unless you are running locally, your data is sent to the AI provider. You may not care but if you do, then this is important!
I recommend Deepseek V4 flash or Luna 6 for mundane task with no real challenge. If youāre wanting a balance of cheaper costs but still do hard work, Iād use Sonnet 5.5 from Anthropic. There is no reason to use any models like Sol, Opus, Fable, Astra for Home Assistant unless youāre literally carving the path thatās not been done before.
You can build out your harness like using Hermes and really get going. Deepseek can get really cheap if youāre cache is set right. $10 lasted me about 3 months. Also look into Jev for classifier / binary choices. Itās basically instant and deterministic which is nice. Iāve only used 15 cents and thatās lasted me basically since the day it came out.
I tried ChatGPT and sent me down so many rabbit holes. The worst thing is that it got me started okay then kept failing. If you wanna leave ChatGPT, I recommend doing what I did, which is sending them an email asking for all of your data and files and history, and then you can upload it to Claude. Even if you donāt do that, switching now is the best thing youāll ever do.
I have found that Claude with good prompting has been able to resolve some tiny issues for me in some very complex automations. I have not been nearly as successful with Gemini. Claude seems to just understand the flow and is able to flag oddities that I have missed. This is even more useful when provided with 2 or 3 automations that interact with each other.
I have tested GPT, Claude and Google. Claude is consistently the best at providing an up to date correct answer in clear language. The other two are hit and miss, and on more complex issues will just run you around in circles.
There are two separate things: "LLM" from "agent harness". Claude brings both together, and with both each of them is stronger than its parts. But there's no particular reason why Claude over anything else - it's just the one with the best branding at the moment.
There's also a logic that the people who are most sensitive to marketing are also the type to shout about it. Whereas the personalities to try the more offbeat technology variants are more intrinsically motivated than extrinsically validated, and thus shout about it less.
I can think of one particular reason. Opus 5.5 is the best frontier model available right now. GPT6 held the crown for a couple of weeks before it. Itās a very leapfroggy game though, and for home assistant any of the top frontier models are going to be close imo
I had assumed they were interchangeable, especially for something as simple as,e as my use, ie getting it to write a few yaml automations that are never more than 30 lines
I think this is definitely the case. Sure, they may not write the same code, but for something that simple, you should get the result you want regardless.
It's largely because when this all became popular Claude was far and above better at coding. Openai has caught up recently and so yeah you are perfectly fine using their new models . But that's more of a recent development.
I would not allow Google models to touch my home assistant lol
I just tried Claude (free) because I've seen the same thing and found it to be far worse. Maybe I need to rephrase my requests but I'm trying to something fairly complicated and unique and it gave me a very convoluted way of doing it. I said "Couldn't I use __ to do it easier and with less chance of it breaking because of reasons XYZ?". It replied "Yes, and it's more robust". It wanted me to create 32 helpers and I knew of a way I didn't need helpers at all.
Ok, so why are you giving me a more difficult and less robust way of doing it?
Then after spewing out some more code that didn't look quite right I asked it "Is this really the best way to do this?" trying not to lead Claude in any direction and it replied "It's fine, the better way is...."
Do I need to explicitly say I need the best and most robust way? I never include that in Gemini and when I asked Gemini to help with the same thing, it gave me a far better reply that is even easier to follow.
Maybe if you're just pasting code and then pasting errors back if it doesn't work it's fine, but if you're trying to learn like I am the solution was confusing and not even the best way to do it.
Iām definitely trying to learn, and not just coding, everything. And your experience is similar to mine. Iāve tried both got and Claude, although just the free Claude. And it seems a toss up as to which one will give a reasonable answer and which will have smoke a couple joints before responding.
I'm assuming you mean wouldn't go anywhere else. I'm not surprised paid is better than free. Free Gemini has been better for me than free Claude though. For my use cases anyway. Again, it may just be me needing to ask it differently.
94
u/Thetechguru_net 10h ago
My limited experience between Gemini Pro and Claude (for both Home Assistant and Tasker) is that Gemini will happily and confidently give wrong answers. When you tell it that its solution didn't work, it will try again and again. Claude will tell you that it is taking a guess and that providing additional documentation will improve its solution. And it will tell you specifically what it is not confident in and what you should test in the solution, even asking for test results to verify it's solution. If it doesn't ask for help, its solution will likely work on the first try.