How to Monitor Cron Jobs with Heartbeats (Dead Man's Switch)

Learn how to monitor cron jobs with heartbeat (dead man's switch) checks: curl in crontab, && chaining, systemd timers, Kubernetes CronJobs and grace periods.

Updated 7 min readBy the Uptime Tracker team

Short answer

To monitor a cron job, give it a unique heartbeat URL and have the job request that URL only after it finishes successfully, for example by appending "&& curl -fsS" followed by the heartbeat URL to the crontab line. The monitoring service expects a ping every interval plus a grace period; if no ping arrives, it alerts you. This catches jobs that crash, hang, never start, or run on a server that is down.

Cron job monitoring works in reverse compared with website monitoring: your job reports in, and silence triggers the alert. This pattern is called heartbeat monitoring, push monitoring or a dead man's switch. It is the only reliable way to notice a scheduled task that silently stopped running, because a job that never starts produces no error to catch.

Why cron jobs fail silently

Cron itself does not tell you when something goes wrong. Common silent failures include:

  • The server or container that runs the schedule is down or was replaced.
  • The crontab was overwritten during a deployment or server rebuild.
  • The script exits with an error, and cron's email output goes to an unread local mailbox.
  • The job hangs on a lock, network call or full disk and never finishes.
  • Environment differences: cron runs with a minimal PATH and no shell profile, so a command that works interactively fails.

Backups, billing runs, report generation, certificate renewals, data syncs and cache warmers are the usual victims. Teams often discover the problem weeks later, when the backup is needed.

How heartbeat monitoring works

  1. You create a heartbeat monitor and receive a unique, secret URL.
  2. You set the expected interval (how often the job runs) and a grace period (how late it may be).
  3. At the end of each successful run, the job sends an HTTP request to the URL.
  4. If no request arrives within interval plus grace, the monitor opens an incident and sends alerts.
  5. When the next ping arrives, the incident resolves.

The grace period absorbs normal variation in run time. A job scheduled hourly that usually takes 4 minutes but sometimes 12 needs a grace period of at least 15 minutes.

Crontab: chain the ping with &&

Use && so the ping is sent only when the job exits successfully. If you used ; instead, the ping would be sent even after a failure, and the monitor would never alert.

# Run every hour; ping only on success
0 * * * * /usr/local/bin/sync-orders.sh && curl -fsS -m 10 --retry 3 https://uptimetracker.live/api/v1/heartbeats/YOUR_TOKEN > /dev/null

What the curl flags do:

  • -f makes curl fail on HTTP errors instead of silently succeeding.
  • -sS hides the progress bar but still prints errors.
  • -m 10 limits the request to 10 seconds so a network problem cannot hang the job.
  • --retry 3 retries transient network failures.

Make sure your script returns a non-zero exit code on failure. In shell scripts, start with set -euo pipefail so any failing command stops the script and propagates the error.

Wrapping multi-step jobs

For jobs with several steps, put the ping at the end of the script rather than in the crontab line:

#!/usr/bin/env bash
set -euo pipefail

pg_dump --format=custom mydb > /backups/mydb.dump
aws s3 cp /backups/mydb.dump s3://example-backups/
rm /backups/mydb.dump

# Only reached if every command above succeeded
curl -fsS -m 10 --retry 3 https://uptimetracker.live/api/v1/heartbeats/YOUR_TOKEN > /dev/null

systemd timers

With systemd, use a oneshot service and add the ping as ExecStartPost, which runs only if ExecStart succeeded:

# /etc/systemd/system/report.service
[Unit]
Description=Generate nightly report

[Service]
Type=oneshot
ExecStart=/usr/local/bin/generate-report
ExecStartPost=/usr/bin/curl -fsS -m 10 --retry 3 https://uptimetracker.live/api/v1/heartbeats/YOUR_TOKEN

# /etc/systemd/system/report.timer
[Timer]
OnCalendar=hourly
Persistent=true

[Install]
WantedBy=timers.target

Kubernetes CronJobs

In Kubernetes, chain the ping in the container command. The image needs curl (or use wget):

apiVersion: batch/v1
kind: CronJob
metadata:
  name: cleanup
spec:
  schedule: "*/15 * * * *"
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      backoffLimit: 2
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: cleanup
              image: example/cleanup:1.4
              command:
                - /bin/sh
                - -c
                - /app/cleanup && curl -fsS -m 10 https://uptimetracker.live/api/v1/heartbeats/YOUR_TOKEN

Heartbeats also catch Kubernetes-specific failures such as a suspended CronJob, missed schedules after a controller outage, or pods stuck pending because the cluster is out of capacity.

Choosing interval and grace period

Job scheduleTypical run timeExpected intervalSuggested grace
Every minuteSeconds1 min2 to 5 min
Every 5 minutesUnder 1 min5 min5 min
Every 15 minutes2 to 5 min15 min10 min
Hourly5 to 20 min1 hour20 to 30 min

Set the grace period to comfortably exceed your slowest normal run, but not so long that a real failure goes unnoticed for hours.

Interval limits in Uptime Tracker

Be aware of the limits before you plan your setup. Uptime Tracker heartbeat monitors accept an expected interval from 15 seconds to 1 hour and a grace period from 0 to 60 minutes (default 5 minutes). Each heartbeat accepts up to 10 signals per minute, either GET or POST. Heartbeats are available on every plan: 5 on Free, 10 on Starter and 50 on Team.

That makes them a good fit for jobs that run every minute up to hourly. Daily and weekly jobs are not directly supported today, because the longest interval plus grace is two hours. If you need to watch a daily job with Uptime Tracker, one workaround is an hourly check script that pings the heartbeat only if the daily job's success marker (for example, a timestamp file it writes on completion) is less than 25 hours old. Otherwise, use a tool that supports long intervals for those jobs.

Only success pings exist; there are no separate start or fail signals. A failure is detected when the expected ping does not arrive, not at the moment the job errors.

Best practices

  • Create one heartbeat per job, named after the job and the server, so an alert tells you exactly what stopped.
  • Treat the heartbeat URL as a secret; anyone with it can send pings.
  • Always use && or a ping at the end of a set -e script so failures are not reported as successes.
  • Add a timeout to the ping so the monitor cannot slow down or hang your job.
  • Route alerts for important jobs, such as backups and billing, to a channel someone checks every day.
  • Test the setup by pausing the job once and confirming the alert arrives.

Learn more on the cron job monitoring feature page.

FAQ

Frequently asked questions

What is a dead man's switch in monitoring?

A dead man's switch is a monitor that alerts when it stops receiving an expected signal. For cron jobs, the job sends an HTTP request after each successful run; if the request does not arrive within the expected interval plus a grace period, the monitor sends an alert.

How do I get notified when a cron job fails?

Append a heartbeat ping to the job with &&, for example: your-command && curl -fsS -m 10 https://your-heartbeat-url. The ping only runs on success, so a failure, hang or missed run results in no ping, and the heartbeat monitor alerts you.

Why use && instead of ; before the curl ping?

With &&, curl runs only if the previous command exited with status 0. With ;, curl runs regardless of the result, so the monitor would receive a ping even when the job failed and would never alert.

How long should the grace period be for a cron job monitor?

Longer than the slowest normal run time of the job plus a small buffer. For an hourly job that usually takes 5 to 15 minutes, a 20 to 30 minute grace period avoids false alerts while still catching failures within the hour.

Can Uptime Tracker monitor daily cron jobs?

Not directly. Uptime Tracker heartbeat intervals range from 15 seconds to 1 hour, with up to 60 minutes of grace. Daily jobs can be covered with an hourly wrapper that pings only if the daily job recently succeeded, or with a tool that supports longer intervals.

Uptime Tracker

Start monitoring in under five minutes

Start on the free plan — commercial use allowed. No credit card, no password, just your email address.

  • Free forever plan
  • No credit card
  • Cancel anytime