← Pulsey Blog

How to Monitor Cron Jobs So You Know When One Stops Running

Cron jobs fail silently. Learn why, how the heartbeat pattern catches missed and failed runs, and how to add it to crontab and GitHub Actions in a few lines.

  • cron monitoring
  • heartbeat monitoring
  • devops
  • how-to

Cron is one of the most reliable tools on a Linux server, and one of the least observable. When a job runs, nothing tells you. When it stops running, nothing tells you either. The first sign is usually a missing backup on the day you need it, or a customer asking why their nightly report never arrived.

This guide covers why cron jobs fail quietly, the common ways people try to watch them, and the pattern that actually catches both kinds of failure: the job that ran and failed, and the job that never ran at all.

Why cron jobs fail silently

Cron’s job is to start a command on a schedule. It does not care what happens after that. A few common ways a job stops doing its work without anyone noticing:

  • The job errors out. A dependency changes, a disk fills up, or a credential expires. The script exits non-zero and cron moves on.
  • The job never starts. The server was replaced and the crontab did not come with it. Someone commented out a line while debugging. The cron daemon is not running in a new container.
  • The environment is different. Cron runs with a minimal PATH and no shell profile, so a command that works in your terminal fails under cron.
  • The job hangs. It starts, blocks on a network call or a lock, and never finishes.
  • Output goes nowhere. Cron can email output through MAILTO, but only if the server has working mail delivery, and most do not.

Checking logs helps only after you already suspect a problem. What you want is to be told.

Approaches that only half work

Emailing output with MAILTO. This reports errors when mail is set up correctly. It says nothing when the job does not run, because no run means no output.

Writing to a log file. Useful for debugging. Nobody reads a log file until something is already wrong.

Alerting from inside the script. Adding a curl to Slack on failure catches crashes the script can see. It misses the case where the script never starts, and it misses hangs.

The gap in all three is the same: they depend on the job running to report that the job did not run.

The heartbeat pattern

The fix is to flip the direction. Instead of the job reporting failure, the job reports success, and something outside the server expects that report.

This is often called a heartbeat or a dead man’s switch:

  1. An outside service gives you a unique URL and a time window, such as “once a day”.
  2. Your job requests that URL when it finishes successfully.
  3. If the request does not arrive within the window, the service alerts you.

Now every failure mode above ends the same way: no ping arrives, and you hear about it. The server can be gone, the crontab can be empty, the job can hang forever. Silence is the signal.

Adding a heartbeat to a crontab entry

The simplest version chains the ping after the job with &&, so it only runs when the job exits successfully:

0 3 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 https://example.com/heartbeat/YOUR_ID

A few details matter here:

  • Use &&, not ;. With ; the ping fires even when the backup fails, which defeats the point.
  • -fsS makes curl fail on HTTP errors, stay quiet on success and still print errors.
  • -m 10 caps the request at 10 seconds so a network issue cannot hang the job.
  • --retry 3 covers a brief network blip between your server and the monitoring service.

If your monitoring service supports explicit failure reports, add a second branch so a failed job alerts right away instead of waiting for the window to expire:

/usr/local/bin/backup.sh \
  && curl -fsS -m 10 https://example.com/heartbeat/YOUR_ID \
  || curl -fsS -m 10 https://example.com/heartbeat/YOUR_ID/fail

Choosing the right window

The window is the one setting that decides whether heartbeats are useful or noisy. Too tight and a job that runs a few minutes late wakes someone up. Too loose and a failure goes unnoticed for days.

A good starting point is about twice the job’s real interval. A job that runs every 5 minutes gets a 10-minute window. A job that runs hourly gets a window a bit longer than an hour. For daily and weekly jobs, leave room for the longest normal run time, not just the start time.

Scheduled jobs outside cron

The same pattern works anywhere a job runs on a schedule:

  • GitHub Actions schedules. Add a final step that pings the heartbeat URL, and use if: always() with the job status so failed runs are reported too.
  • Laravel’s scheduler, Celery beat, systemd timers or Kubernetes CronJobs. Ping at the end of the task, or use the framework’s success and failure hooks.
  • Long-running workers. Have the worker ping on a regular interval, such as every few minutes, so a crashed or stuck worker goes quiet.

Monitoring cron jobs with Pulsey

Pulsey, our uptime monitoring service, has heartbeats built in. You create a heartbeat, choose a timeframe from 10 minutes to 1 month, and Pulsey gives you a unique ping URL. Its page has ready-made snippets for curl, cron, Laravel, PHP, Python and Node, plus a Send Test Ping button to confirm it works.

Pulsey evaluates heartbeats every minute on every plan. When a ping is late, the heartbeat is marked missing, an incident opens, and the alert goes to email, Slack, Microsoft Teams, Discord, Telegram or a webhook, with SMS, PagerDuty and Opsgenie on paid plans. Requesting the URL with /fail on the end marks the job failed immediately and can include a short message in the alert. When the next normal ping arrives, the heartbeat returns to beating and the incident resolves on its own.

For GitHub Actions, add a last step that runs even when earlier steps fail and pings the right URL for the job’s result. Store the heartbeat URL as a repository secret:

- name: Report to Pulsey
  if: always()
  env:
    PULSEY_URL: ${{ secrets.PULSEY_HEARTBEAT_URL }}
    RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
  run: |
    if [ "${{ job.status }}" = "success" ]; then
      curl -fsS -m 10 --retry 3 "$PULSEY_URL"
    else
      curl -fsS -m 10 --retry 3 --data-urlencode "message=Workflow ${{ job.status }}: $RUN_URL" "$PULSEY_URL/fail"
    fi

A failed run then alerts right away, with a link to the run in the alert.

The free plan includes 50 heartbeats alongside 50 website and API monitors, with no credit card required.

Start watching your jobs

Pick the one scheduled job you would least like to discover was broken, usually the backup, and put a heartbeat on it today. It takes one line in your crontab.

Create a free Pulsey account and add your first heartbeat in about five minutes.

Find out before your customers do.

Add an endpoint, and get paged within minutes when it breaks.

No credit card · 5-minute setup · 50 monitors free forever