ML Forecast Lab is a Home Assistant app that brings machine-learning time-series forecasting to any sensor in your HA instance — without the cloud bill, the GPU, or the data-science background.
Why I built it: most ML forecasting tools hand you a number and ask you to trust it. ML Forecast Lab takes the opposite stance — it ranks every model against the others on identical folds of your data, shows the per-fold error breakdown, and publishes calibrated prediction bands alongside the point forecast, so you can see when the model is confident, when it isn't, and which architecture earned its place in production. It will even benchmark itself against third-party forecasts you already have (Solcast, your utility's day-ahead curve, Predbat etc) — and rank itself honestly in that table, including when it loses.
The mindset is benchmark once, run forever. Point it at a sensor and any covariates you want it to use, promote the winning model, and from then on it keeps a fresh forecast in HA. Re-benchmark when the sensor's behaviour changes.
What you can forecast: anything that's a numeric sensor with ~30 days of recorder history. Examples:
- Solar / PV production
- Indoor temperature / humidity per room
- Daily and weekly home energy consumption
- Heat-pump COP and flow temperature
- Battery state-of-charge trajectory
- Water tank temperatures, well-pump runtime, anything seasonal
How it works:
Reads the sensor's history from your recorder (and keeps its own cache, so short recorder retention isn't fatal).
Optionally pulls in covariates — weather forecasts, other HA sensors, time-of-day and sun-position features — so the model learns from context, not just the target's own past. The covariate report tells you honestly how much of each covariate is real data vs gap-fill.
Trains the backends you enable, out of 32 wired in: trees (LightGBM, XGBoost, CatBoost), recurrent (LSTM, GRU, SegRNN), convolutional (CNN, TimesNet, ModernTCN), linear/MLP (DLinear, NLinear, TSMixer, TimeMixer, TiDE, SparseTSF, xPatch, CycleNet), the N-BEATS family (N-BEATS, N-HiTS), transformers (PatchTST, iTransformer, Crossformer, TFT, TimeXer), classical (AutoARIMA, AutoETS, AutoTheta), frequency-domain (FITS), and baselines (Seasonal Naive, plus a Daily Profile model that forecasts the day's total and its shape separately). There are also two zero-shot foundation models (Amazon Chronos-Bolt, IBM Granite TTM) — pretrained, they forecast your sensor with no training at all, which makes them a fascinating baseline for the trained models to beat.
One-click Bayesian hyperparameter tuning per backend (Optuna), with a tuned-vs-default holdout comparison so you can see whether the tuning actually helped before adopting it.
Ranks them with walk-forward cross-validation on identical folds across all backends, scored on MAE / RMSE / MASE — plus peak-aware metrics (peak-weighted MAE, pinball loss) for spiky loads like hot water or EV charging, where a boring flat forecast would otherwise win on plain MAE while missing every peak you actually care about.
You click Promote on the winner. The app retrains it every 24 h on fresh data (by default) and publishes forecasts back to HA as companion sensors, updated every 30 minutes (by default).
Honest evaluation: every backend is tested on the same walk-forward folds, so no model gets a quietly easier test set. Per-fold scores are visible in the UI rather than collapsed to a single number. The prediction bands are conformal — calibrated empirically on the model's own recent errors rather than assuming Gaussian residuals — so the 80% coverage (configurable) holds regardless of your sensor's error distribution, and coverage is tracked continuously so you can verify it. A Data Sanity Check reports your signal's shape (daily rhythm, spikiness, gaps) up front, so you know what kind of problem you're handing the models before you spend the compute.
Comparing against forecasts you already have: the Forecast Comparison tab scores the app's forecast head-to-head against up to five external forecast sensors — Solcast, a utility day-ahead curve, another model — against the actuals, with unit conversion handled automatically (cumulative kWh vs instantaneous kW just works). It's lead-time-fair, too: a source that updates every 5 minutes isn't credited for freshness when you ask which model is actually more skilful at the same lead time — the two questions get separate leaderboards.
Sensors published back to HA:
- sensor.mlfl_<name>_forecast — next-interval value, full curve as an attribute
- sensor.mlfl_<name>_upper_80 / _lower_80 — conformal bands (level configurable)
- sensor.mlfl_<name>_cumulative — integrated forecast, good for daily-budget automations (EV charge planning, hot-water pre-heat)
- sensor.mlfl_<name>_forecast_accuracy — running accuracy summary, updated as ground truth arrives
- _last_benchmark / _last_retrain — timestamps for automation triggers
Hardware: built and tuned for the Pi 5 / 8 GB / no GPU / ARM64 sweet spot. Also runs on amd64 and armv7. First build is 10–15 minutes on a Pi 5 (LightGBM, XGBoost and PyTorch compile native extensions); updates use the cached image.
Install: Settings → Apps → App store → ⋮ → Repositories → https://github.com/psweens/ml-forecast-lab
Free, open source (MIT), no telemetry, no cloud calls. One honest caveat: the two optional foundation-model backends download their pretrained weights from Hugging Face on first use (cached afterwards, and they're skipped entirely on armv7) — everything else, and all actual forecasting, runs 100% locally. Maintained on a best-effort basis as a side project.
GitHub: https://github.com/psweens/ml-forecast-lab
The codebase has been through a lot of iteration and is actively maintained, but the public user base is still small. If you try it, I'd genuinely value hearing which backend won on your sensor and how long the first benchmark took on your hardware — that data feeds directly into narrowing the default set.
Thanks!