How to monitor a cron job (and know when it silently stops)

Cron does not tell you when a job fails or never runs. Add a dead man's switch check with one curl call and get alerted when the ping goes missing.

To monitor a cron job, have the job call a URL every time it finishes successfully, and let an external service alert you when that call does not arrive on time. This is a dead man's switch: silence is the failure signal. With FlatLyne you create a check with the job's schedule and a grace period, then add one curl to the end of the crontab line.

Why cron fails silently

Cron starts a command and forgets about it. It has no idea whether the job did its work, so these failures all look the same from the outside: nothing happens.

  • No mail. Cron reports output by local mail. Many servers have no mail transfer agent, or MAILTO is empty, so errors go nowhere. Even with mail working, a job that never started produces no output to send.
  • Environment differences. Cron runs with a minimal PATH and none of your shell profile. A script that works in your terminal can fail under cron because node, pg_dump or an environment variable is missing.
  • The server is down. If the machine is off, or crond is not running, nothing runs and nothing complains. Only an outside observer notices.
  • Overlapping runs. A run that takes longer than the interval collides with the next one. Both may fight over a lock or a file, or the second one may exit early without doing the work.
  • DST and timezone. Cron runs in the server's timezone. Around a daylight saving change, a job scheduled in the skipped hour does not run that day, and one in the repeated hour may run twice.

A check on an external service covers all five, because it only cares about one thing: did a successful ping arrive on time?

Step 1: Create a check in FlatLyne

Sign in, open Checks and select Create check. Set:

  • Schedule type: Cron expression, and paste the same expression as your crontab line, for example 0 2 * * * (02:00 every day). Pick the Timezone your cron daemon uses.
  • Grace period: how long FlatLyne waits after the due time before marking the check Down.

For a cron check, the due time is the scheduled start time, not the end of the run. A job that pings when it finishes sits in Grace for its whole run, every time, without alerting. So make the grace period longer than your job's longest normal run. A backup that takes 5 to 25 minutes is fine with a 1 hour grace period.

See the quickstart for the full walkthrough and schedules for how FlatLyne computes due times. The Free plan includes 10 checks with a shortest schedule of 1 minute.

Copy the check's ping URL from the Ping endpoint field. It looks like this:

Ping URL
https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NI

Treat it like a password. Anyone with the URL can mark your check up or down.

Step 2: Add the ping to your crontab

The simplest version pings only when the job succeeds. && runs curl only if the job exits with 0:

crontab
PING_URL=https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NI
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL"

The flags matter: -f fails on HTTP errors, -sS hides progress but keeps errors, -m 10 stops a stalled request after 10 seconds, --retry 5 retries timeouts and temporary server errors, and -o /dev/null discards the PONG reply so cron does not mail it to you. If the job fails, no ping arrives and FlatLyne alerts once the due time plus grace has passed.

Report failures immediately and time the run

To hear about a failure right away instead of after the grace period, send a fail ping. To measure duration, send a start ping first. A short wrapper does both and keeps the job's exit status:

/usr/local/bin/backup-with-pings.sh
#!/bin/bash
PING_URL="https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NI"

fl_ping() { curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL$1" || true; }

fl_ping /start
if /usr/local/bin/backup.sh; then
  fl_ping ""
else
  status=$?
  fl_ping /fail
  exit "$status"
fi
crontab
0 2 * * * /usr/local/bin/backup-with-pings.sh

A /fail ping marks the check Down at once and sends an alert. A /start ping records the start without changing the check's state or due time, so a job that starts and then hangs still goes Down on schedule. The ping after a start ping shows the run's duration on the check's page.

If you prefer one line, send the exit code: /usr/local/bin/backup.sh; curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL/$?". Code 0 counts as success and any other number as a failure.

Each check accepts 3 pings per minute (extra pings get 429), and a start plus a final ping uses 2. See ping URLs and reliability.

Ping by slug with a ping key

Instead of one URL per check, a project ping key lets you ping by name: https://api.flatlyne.com/<ping-key>/<slug>. Give the check the slug backup-db-nightly, then:

crontab
PING_URL=https://api.flatlyne.com/<ping-key>/backup-db-nightly
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL"

/start, /fail and /<exit-code> work the same way after the slug. Ping keys can be revoked and replaced, but a leaked key reaches every slugged check in the project. See ping keys.

Step 3: Set up alerts

FlatLyne alerts when a check goes Down and again when it recovers. By default it emails every member of the project. You can add Slack, PagerDuty and webhook channels from the dashboard, and the Free plan allows 2 alert channels. See alerts.

Common pitfalls

  • Relative paths and missing env vars. Use absolute paths in crontab lines and scripts, and set anything the job needs in the script itself.
  • A && B || C on one line. If A succeeds but the success ping B fails on the network, C runs and sends a false fail ping. Use the if/else wrapper above.
  • Unescaped % in crontab. Cron turns % into a newline. Escape it as \% or move the command into a script.
  • Grace period shorter than the run. The check sits in Grace, and a slow run turns into a false Down.
  • Overlapping runs. FlatLyne does not give runs an ID, so overlapping runs mix in the ping history. Prevent them with flock -n /tmp/backup.lock /usr/local/bin/backup.sh.
  • DST. Run the server and the check in UTC, or use a cron check with the same expression and timezone, and avoid 01:00 to 02:59 local time.

FAQ

What is the difference between a start ping and a success ping?

A success ping says the job finished fine and sets the next due time. A start ping only records that a run began, and leaves the status and due time alone. Use both to see how long each run takes.

What happens if the server is down and cron never runs?

No ping arrives, so the due time and the grace period pass and the check goes Down. That is the main reason to monitor from outside the machine.

Will a failed ping request break my job?

Not if you use the wrapper above. || true ignores the ping failure, and a missing ping only causes an alert, never a silent miss.

Know when your jobs stop running.

Try FlatLyne free