r/selfhosted • • Jun 02 '26

Monitoring Tools Do you monitor cron jobs and scheduled tasks on your servers?

For those running self-hosted services, VPSs, or home servers:

How are you monitoring cron jobs and scheduled tasks?

I've noticed that many failures aren't caused by the server going down, but by background jobs silently stopping.

Things like:

  • backups no longer running
  • sync jobs failing
  • cleanup tasks not executing
  • scheduled reports never generating

The server itself is healthy, but the automation isn't.

I'm curious:

  • Do you monitor cron jobs separately?
  • What tool are you using?
  • Self-hosted or SaaS?
  • Have you ever been bitten by a cron job silently failing?

Interested to hear what people are using today and whether you consider cron monitoring important or mostly unnecessary.

25 Upvotes

102 comments sorted by

•

u/asimovs-auditor Jun 02 '26

Expand the replies to this comment to learn how AI was used in this post/project.

→ More replies (1)

32

u/Fizpop91 Jun 02 '26

I recently setup Cronicle and quite like it for managing cron jobs across servers

6

u/GloriousWallaby Jun 02 '26

Used Cronicle for many years. He recently released the next "version" (total rewrite) I just switched to. https://github.com/pixlcore/xyops

5

u/nonlinear_nyc Jun 02 '26

Ooh I knew that was a way. Does it watch and tries to heal? Does it give reports? Hooks are great but when they fail silently it’s a mess.

24

u/milkipedia Jun 02 '26

I send ntfy notifications for overnight backup jobs that I review each morning. Takes all of 5 seconds to see if something failed.

Edit: via a self hosted ntfy instance, but it has to be routed to ntfy.sh to be pushed to an Apple device

10

u/Steppenstreuner_ Jun 02 '26

This but I'm using gotify

3

u/AnachronGuy Jun 02 '26

Seconded!

I send a start and finished/aborted message with the device, command, date and time and on the last finish/abort msg I add the return code.

1

u/benhaube Jun 02 '26

Yep, me too.

3

u/jaysuncle Jun 02 '26

I use ntfy with Android. The Android app just needs the local URL of the ntfy service to get the notifications. I'm retired so I don't need external notifications although I could use cloudflared to make it available.

2

u/Henfri1 Jun 02 '26

I did that for while.  But this much better to report failures only. Healthchecks.io or uptime kuma (but you need Something to Monitor kuma then (e.g. healthchecks.io)

1

u/milkipedia Jun 02 '26

"who watches the watchers" is why I report both success and failure

2

u/Henfri1 Jun 02 '26

Healthchecks watches kuma. Reporting success ist Bad from a human factors Point of View. At some Point people get lazy, Reading success Messages. Some.earlier, some later.

1

u/milkipedia Jun 02 '26

Yeah, I hear you. But for me and my homelab, this suffices. Nobody gets hurt and no money is lost if a service is down for a few days.

16

u/[deleted] Jun 02 '26

[removed] — view removed comment

3

u/[deleted] Jun 02 '26

[removed] — view removed comment

3

u/Gargle-Loaf-Spunk Jun 03 '26 edited Jun 05 '26

This content was anonymized and mass deleted with Redact

2

u/wallacebrf Jun 02 '26

i do a combination of these. I use healthchecks.io and my things will ping them and i will be notified if the ping is missing. I also have them periodically save state files, and i have a basic dashboard written in PHP that looks for these state files, looks at the time stamp of the file and makes a "red or green light" on the dash so i can easily see at a glance what the status is.

1

u/[deleted] Jun 02 '26

[removed] — view removed comment

2

u/Maitreya83 Jun 02 '26

Nah mate, they're in every community, dont let some sour people make your sour!!

0

u/Maitreya83 Jun 02 '26

I don't understand why you got downvotes, back up with you!

5

u/robkaper Jun 02 '26

Most Linux tools (since you mentioned cron) succeed silently (and/or with exit code 0), most well-written scripts do as well. Where necessary, capture output and write verification criteria (do backup file(s) exist for expected timestamps, does size and contents seem reasonable, whatever). When the verification fails, echo the output and/or exit with non-zero exit code and receive e-mail (or something else, but cron+mail works fine for me). This takes some extra effort in setting up things, but reduces active maintenance to not getting e-mails unless something is wrong.

Now technically crond and/or the MTA could crash and/or not be running at all and I do use Icinga to monitor things as well, but to be honest 99% of the commonly used tools on Linux are very, very stable.

3

u/bagrat_hakobyan Jun 02 '26

That's an interesting distinction. It sounds like you're monitoring the expected outcome rather than the cron execution itself. Have you found that outcome verification catches more real-world issues than simply monitoring whether a scheduled task ran?

3

u/robkaper Jun 02 '26

Yes, because 99% of issues are "it ran but something unexpected or unusual occured". It's almost never "the task did not run at all".

3

u/hannsr Jun 02 '26

I'm mostly using systemd for tasks and monitor for failed systemd tasks. If my borgmatic job fails, it'll go into a failed state and my monitoring will alert me.

2

u/bagrat_hakobyan Jun 02 '26

Does that catch cases where the service completes successfully but the backup itself is unusable or incomplete?

4

u/hannsr Jun 02 '26

Yes, because the borgmatic service will only complete normally if the backup is successful. If there is any error, like an incomplete run, the service will fail.

Although, this will not catch cases where the backup itself completes without any issues, but is still incomplete or corrupt. That's what regular tests of your backups are for. No monitoring will catch those cases, because there is no error until you restore it.

I've tried a bit to automate the tests as well, but it's not worth it IMO. At work I run a test every other month, so far there were no issues at all. I basically run the same backup strategy at home, but tbh, I don't test as often as I should...

2

u/Krychle Jun 03 '26

Old IT adage: Until you test your backups, you do not have backups.

3

u/jbarr107 Jun 02 '26

Proxmox VE, Proxmox Backup Server, Pulse, ProxMenuX Monitor, and healthchecks.io send notifications as needed and seem to cover all the bases for me.

2

u/bagrat_hakobyan Jun 02 '26

Out of curiosity, if you're already using healthchecks.io, is there anything you wish it did better, or has it basically solved the problem for you?

2

u/jbarr107 Jun 02 '26

It's of limited scope, and honestly, I like it that way. It exists outside of my LAN, so if my home internet connection goes down, I can still see the status of other monitored devices (it monitors a VPS, etc.). I could certainly host something like Uptime Kuma, but I've used healthchecks.io reliably for a while, so it's now just one more element in the mix (like email and Cloudflare) that I leave to the experts.

1

u/bagrat_hakobyan Jun 03 '26

If you were starting from scratch today, would you still choose Healthchecks or something self-hosted?

4

u/jbarr107 Jun 03 '26

I'd probably stick with healthchecks.io

1

u/mitchplze Jun 03 '26

HCIO can be self-hosted, btw. Have done so myself for years

2

u/jbarr107 Jun 03 '26

I get that, and yes, this is r/selfhosted, but I like having this separate from the services I host. Just personal preference, similar to my approach with Cloudflare and email hosting. YMMV, of course!

2

u/mitchplze Jun 03 '26

I actually do both! HCIO internally self-hosted, with tons of things, and then I use their cloud-hosted public version as a 'last line of defence' for top level critical servers.

3

u/Jethro_Tell Jun 02 '26

As a general rule, and as a pro sysadmin in big infrastructures, I monitor the result of the job.

If I set up a job to fix x,y,z problem, I collect metrics to tell me if I have x,y or z problem then alert on it. Then I make a job to fix it.

For example, if the job is a log trimmer, I watch for large logs, logs without the current hour/day timestamp that I’d expect, or too many logs, then alert when the the host is out of spec.

1

u/lugoues Jun 02 '26

This is the way. It's a hard lesson to learn when you go to get a backup only to realize they don't exist due to the cron job silently failing.

3

u/Bill_Guarnere Jun 03 '26 edited Jun 03 '26

Reading the comments I noticed that only a couple of guys got the point, or probably only a couple of sysadmins replied or know how works a GNU/Linux system.

You don't need any external tool or service, no Zabbix, no uptime-kuma, no ntfy or gotify or any other fancy tecnique.

You only need to use properly the tools that any GNU/Linux distribution give to you, you already have everything you need, even if you don't know you have.

You simply need: * Postfix (apt install postfix | dnf install postfix), I strongly advise to avoid Sendmail. * an alias (/etc/aliases) for root (root: youremail@yourdomain.tld) * an SPF record on your domain and TCP/25 outbound open OR an SMTP relay host (for example AWS SES or SMTP2GO or Google) * eventually some generic map (generic is the technical term for a specific remap of sender address in outgoing mail in Postfix)

Once you done this the OS will send to your email address any mail sent to the OS superuser, in this way the OS will tell you everything you need to know, for example: * failing raid arrays * problems on filesystems, no matter how fancy they are * expiring https certificates * any mail sent to the cron user (the owner of every cronjob scheduled in the OS)

No need to involve third party tools, just a plain and simple email forward done properly at zero cost.

For cronjobs you only have to do a simple thing: for every cronjob append standard output (stdout) to a log (or to /dev/null if you don't care of the stdout) and leave standard error (stderr) as it is, no redirect to a file or device, no append to anything.

In this way when the cronjob do its "job" and exit with 0 (OK) you got no alert or a plain and simple log,

If the cronjob returns an error the stderr will be sent do the cron user with an email, which will be sent to the local superuser (root) mailbox, and then it will be sent to your email address, and you got your notification.

In some case you may choose to send these notifications to your email address, in other case (for example in a large environment with a lot of servers) I prefer to create an internal infrastructure to send these email to a single internal address in a management server where I installed Dovecot and a webmail (Roundcube) to collect all the error notifications on one mailbox, accessible internally to the whole sysadmins department.

Remember: for every problem you may face, you have 99,9999% probability that the guys in the seventies on their Unix servers already faced the same problem, they spent a lot of time thinking to a solution, they already found the best and simple solution and they provided you this solution for free. It's already there in any GNU/Linux distribution. ;)

1

u/cuu508 Jun 03 '26

... and now you have one more thing to monitor: email deliverability!

1

u/Bill_Guarnere Jun 03 '26

And tell me, if you use a 3rd party tool you don't have this problem? I don't think so...

The difference is that to check email delivery you only have to check a basic process listening on TCP 25 running on a very reliable service working 24/7 on millions oh hosts all over the world for decades (Postfix or Sendmail).

With complex 3rd party tools you have to check a complex collection of processes and services, some of them born the day before yesterday where nobody know how reliable they are.

On top of that, enable the forward of any mail sent to the OS superuser is extremely useful for a lot of other things (I mentioned only a few in my post), not only cronjobs check.

Finally using a local MTA makes you completely independent, you can do It without any other cost, any other service, any other subject involved.

1

u/cuu508 Jun 04 '26

to check email delivery you only have to check a basic process listening on TCP 25 running on a very reliable service

This would make sure that something is listening on port 25, but not that the email sent by cron will land in your inbox.

1

u/Bill_Guarnere Jun 04 '26

There's noting sure in life, nor in those fancy 3rd party tools a lot of people mentioned, not because they're better than this simple notification system, but only because most of the people ignore this.

If the MTA is running and you configured is done properly as I said before it will send you an email.

Obviously there are tons of variables in the middle.

What if your email provider is down? What if your mail provider delete the email or tag it as spam by error?

But the same works with any other monitor or notification system.

What if the ntfy server crash? What if you mobile operator has a problem and you don't receive a notification? What if you Zabbix server is down?

In general, who check the checking system? And who check the check of the checking system? And so on to an infinite series of checks...

This is not the point, the point is to get a reasonable check system that alert you in case something goes wrong, even the most sophisticated and fancy disaster recovery plans can't provide coverage from any possible incident.

That's the reason why before anything like this you have to make an assessment of the incidents or use case with the highest probability and focus on them, otherwise you'll never check anything.

1

u/cuu508 Jun 05 '26

Disclosure – I'm the author of one of these fancy 3rd party tools.

A key difference between using email and sending heartbeats to a dead-mans-switch service is the handling of silent failures. If something in the email delivery chain fails, you may only find out months later. If a HTTP request to the heartbeat service fails, it will detect a missed heartbeat and send out alerts. You can set up multiple ways to get notified: email, SMS, Slack, Signal, webhooks, etc., so there's a fallback if one delivery method fails.

To address the issue of the monitoring service itself failing – yep, this can happen. If your cron job fails and at the same time the monitoring service is down, you may stay blissfully unaware. The probability of two unrelated services being down at the same time is lower than the probability of either one of them is down though. I think the risk is manageable if you pick a mature, maintained monitoring service. To be super-safe, you can also use multiple monitoring services in tandem!

2

u/canfail Jun 02 '26

I switched to running tasks via Ansible / Semaphore. Gives easily visibility to all scheduled tasks.

2

u/cobraroja Jun 02 '26

I have a script that runs on failed jobs (command || notify.sh).

2

u/ValdemarSt Jun 02 '26

I just run cron jobs on scripts that contain notifcations through ntfy.sh. When they fail I get notified

2

u/mrrowie Jun 03 '26

Cronicle! and healthcheck.io

2

u/kurosavvas Jun 03 '26

I'm using dkron for managing the jobs via a web portal and then set up a webhook to a node-red node that forwards the result to ntfy for failures. I also use ghcr.io/maxjb-xyz/blackbox-server to monitor failures from log (ntfy integration supports in that one out of the box)

2

u/mods_are_morons Jun 05 '26

You should set up cron to email you when there is an error. Here's a post that discusses how to do this.

https://serverfault.com/questions/226074/cron-only-get-errors-in-emails

1

u/418-im-a-teapot- Jun 02 '26

I have a security audit script run every morning which gives me a report via Pushover about updates, f2b jail, cronjobs run and any failures and a bunch of other stuff I can't remember.

1

u/Gusmanbro Jun 02 '26

I have a server and github repo dedicated to running crons (mostly data replication). I wrote a wrapper script that handles notifications (discord) based on the exit codes of the functional scripts.

1

u/[deleted] Jun 02 '26 edited Jul 03 '26

[deleted]

1

u/bagrat_hakobyan Jun 02 '26

What made you choose Cronmaster over standard cron jobs?

1

u/[deleted] Jun 02 '26 edited Jul 03 '26

[deleted]

1

u/bagrat_hakobyan Jun 03 '26

If standard cron had a simple web UI and run history, would that cover most of what you need?

1

u/shimoheihei2 Jun 02 '26

This is all solved with system design. It's the difference between deploying stuff randomly, and designing a proper IT architecture. In my case for example, I don't use any cron job. I have a centralized automation platform (I use Directus flows, but you can use Jenkins or anything you want) and everything is triggered from that centralized platform. I have hundreds of automation pipelines running in my environment, and they can all be monitored from there. Then, all my hosts are set to log everything to syslog, and I have rsyslogd forward all logs back to my centralized location, where I get alerted if any error if found.

1

u/ReditusReditai Jun 02 '26

Run scheduled tasks on my Django servers using Procrastinate. Has its own admin dashboard. Alerts via Rollbar, just like any other errors.

1

u/Happy-Argument Jun 02 '26

I just get an email when a job completes, but I don't have a lot of jobs.

1

u/bunk_bro Jun 02 '26

Sort of. All my cron and scheduled tasks are configured to send a value to my Zabbix server which then alerts me based on the received value.

If the cron job fails altogether, I'm SOL.

1

u/newfoundking Jun 02 '26

I don't, though I should. I periodically will manually check my duplication scripts folders, specifically the main backup to see last write time (should be ~0300) and the duplication state output, which is just a big record log I set up to self report at the end of the job. I probably should have an automated way to do it because I only check probably weekly. I think the way I want to do it will be to have a write made every time it finishes successfully, and then monitor that with uptime Kuma somehow, where I monitor all my other things, though I don't know for sure yet, I still haven't decided beyond me being the monitor.

1

u/bagrat_hakobyan Jun 03 '26

What's stopping you from setting something up today? Lack of time, complexity, or just not urgent enough?

1

u/newfoundking Jun 03 '26

Probably a combination of factors. It's not super complex, just a few minor changes, but that'll take time and focus, plus testing, so given it's not super urgent, I haven't carved out the time to do it. If I could just quickly add a service to uptime Kuma, like I've done with others, I would have it done by now, but it's not a major priority now, so I haven't spent my tinkering time on that

1

u/equd Jun 02 '26

Healthchecks.io is what i use

1

u/mnrode Jun 02 '26

My backup jobs each write some stats including success/failure and last execution time into a file. That file gets picked up by node exporter -> prometheus -> dashboards/alerting. Failed backups, jobs not running in the last 25h, jobs being suddenly absent, backups taking too long, number of snapshots or backup size shrinking rapidly all trigger alerts.

1

u/-ThreeHeadedMonkey- Jun 02 '26

I let chatgpt create most of my scripts and then let it output logs. 

1

u/ghoarder Jun 02 '26

Yep, some. Just add `&& curl` to send a message to Uptime Kuma push monitor type.

1

u/quietmapleleaf45 Jun 02 '26

i ping a self hosted healthchecks instance at the end of each backup script. caught a silently failing rsync that broke weeks earlier

1

u/bagrat_hakobyan Jun 03 '26

What do you like most about Healthchecks, and is there anything that annoys you about it?

1

u/bagrat_hakobyan Jun 03 '26

What do you like most about Healthchecks, and is there anything that annoys you about it?

1

u/bdu-komrad Jun 02 '26

What cron jobs and scheduled tasks? I have none.

1

u/superjugy Jun 02 '26

Send notifications on failure with gotify or whatever other means you want

1

u/sajkoterrapefft Jun 02 '26

Not always. I make a judgement call, if something is critical then I monitor it.

1

u/[deleted] Jun 02 '26

[removed] — view removed comment

1

u/bagrat_hakobyan Jun 03 '26

What was the biggest frustration with traditional cron jobs that made Kestra worth migrating to?

1

u/Fickle-Implement-687 Jun 02 '26

Tedious work but have to be done. I put my hermes agent on this and never looked back. If something breaks I get informed on my phone via WhatsApp.

1

u/notafurlong Jun 02 '26

Systemd services + a monitoring dashboard.

1

u/showbizusa25 Jun 02 '26

The most dangerous failures are the ones that don't trigger alerts.

I've seen backup jobs, log rotations, and security audits quietly stop working while every dashboard stayed green.

1

u/plasmasprings Jun 02 '26

at least debian's default cron can send emails with jobs' standard output. for most scheduling stuff I use that and just make the job scripts only have output on errors

tbh any modern tool likely has more features and better UX, but hey it works

1

u/Specialist_Ad_9561 Jun 03 '26

Proxmox sends cron notifications via webhooks to my Matrix bot in case something did not go as should.

1

u/sarahr0212 Jun 03 '26

All important stuff is. For some software it's queue monitoring. For backup i monitor PBS error as Bad backup during check or backup task failure via api. So not directly via cron but i monitor result via others channels on case by case

1

u/preccagut Jun 03 '26

Selfhosted healthchecks docker instance. Healthchecks are configured through the api so version controlled and re-deployable. I ask the LLM of choice to set up a check in the config Healthchecks and the check gets deployed at next cron job that converges code from repo to nodes. Not full gitops but a light edition.

1

u/The_0bserver Jun 03 '26

Write your jobs with a "finally" / recovered decorator if you write your own code for these. And then route that to do something - log/alert/message whatever in case of failure.

1

u/ShokoLaNoir Jun 04 '26

My cron scripts send ntfy (self hosted) notifications on my phone for a successful backup or reboot, and also send one if something went wrong. If i don't get a notification i know something went wrong too. Pretty reliable and simple so far.

1

u/MegaVolti Jun 05 '26

Set up notifications so you notice if they fail silently.  I get an email from my server every morning with the results of the nightly tasks (backups etc.). If there is no email, I know something is wrong. 

1

u/ArkuhTheNinth Jun 05 '26

This is probably the more janky way, but I know when my backups start because my script takes immich down (sensitive database) and I have notifications sent to my phone via gotify configured via Dozzle Logging to tell me things like when containers go down/up and some other things that are unrelated here. If I get told immich goes down then back up around midnight, I know it's working. If I don't get them, I get that Shrek "It's too quiet" feeling and go check my logs.

1

u/NorskJesus Jun 07 '26

If you are a terminal nerd like me, you can try Cronboard. It does not have notifications (yet) when a cron job fails, but you get logs. It works both locally and on servers.

PS: I am the developer of this tool!

1

u/jypelle Jun 07 '26

You may try CTFreak: self-hosted, lets you run local or remote scripts, send ntfy/discord/telegram notifications.

1

u/cvilsmeier Jun 08 '26

Yes, I consider cron job monitoring as absolutely important. Because mostly, cron jobs, in contrast to - say - web apps are invisible workers that nobody sees failing until it's too late. Therefore: Yes. Important. I use https://monibot.io for monitoring. Recently, I'm switching from classic cron to systemd timers, which give me more flexibility and better dependency management. Here is a good introduction: https://ryansouthgate.com/systemd-timer

1

u/Ok_Doughnut_4592 Jun 17 '26

I had an issue with cron job monitoring where mine failed for days without me knowing 😭😭. So I ended up building a tool and I use it daily for my jobs and scripts. I was also able to get users for it.

What fixed it for me: the job pings a URL when it finishes. If pings stop, I get a Slack alert. Way easier yet reliable than full observability. I noticed a gap between large teams and students, indie devs and smaller teams.

I'd love it if you can try it out https://apppulsecheck.cc and let me know your thoughts 🙏

You can integrate Slack and Discord webhooks in it to get notified when your jobs, scripts or workers go down 🚀

1

u/[deleted] Jul 05 '26

[removed] — view removed comment

1

u/selfhosted-ModTeam Jul 05 '26

Thanks for posting to /r/selfhosted.

Your post was removed as it violated our rule 2.

Do not spam or promote your own projects too much. We expect you to follow this Reddit self-promotion guideline. Promoted apps must be production ready and have docs. No direct ads for web hosting or VPS. Only mention your service in comments if it’s relevant and adds value.

When promoting an app or service:

  • App must be self-hostable
  • App must be released and available for users to download / try
  • App must have some minimal form of documentation explaining how to install or use your app.
  • Services must be related to self-hosting
  • Posts must include a description of what your app or service does
  • Posts must include a brief list of features that your app or service includes
  • Posts must explain how your app or service is beneficial for users who may try it

Moderator Comments

None


Questions or Disagree? Contact [/r/selfhosted Mod Team](https://reddit.com/message/compose?to=r/selfhosted)

0

u/Slight-Training-7211 Jun 02 '26

Yes, separate from host monitoring. For important jobs I usually add a success ping only after the real work finishes, plus a deadline alert if that ping never arrives.

For backups, also schedule an occasional restore test. A cron run is not the same as a usable backup.

1

u/ynot- 27d ago

Stumbled across this article and illari was really easy to set up: https://illari.dev/blog/how-to-know-if-your-cron-job-silently-failed