How to monitor a cron job (and know when it silently stops)
Cron does not tell you when a job fails or never runs. Add a dead man's switch check with one curl call and get alerted when the ping goes missing.
To monitor a cron job, have the job call a URL every time it finishes successfully, and let an external service alert you when that call does not arrive on time. This is a dead man's switch: silence is the failure signal. With FlatLyne you create a check with the job's schedule and a grace period, then add one curl to the end of the crontab line.
Why cron fails silently
Cron starts a command and forgets about it. It has no idea whether the job did its work, so these failures all look the same from the outside: nothing happens.
- No mail. Cron reports output by local mail. Many servers have no mail transfer agent, or
MAILTOis empty, so errors go nowhere. Even with mail working, a job that never started produces no output to send. - Environment differences. Cron runs with a minimal
PATHand none of your shell profile. A script that works in your terminal can fail under cron becausenode,pg_dumpor an environment variable is missing. - The server is down. If the machine is off, or crond is not running, nothing runs and nothing complains. Only an outside observer notices.
- Overlapping runs. A run that takes longer than the interval collides with the next one. Both may fight over a lock or a file, or the second one may exit early without doing the work.
- DST and timezone. Cron runs in the server's timezone. Around a daylight saving change, a job scheduled in the skipped hour does not run that day, and one in the repeated hour may run twice.
A check on an external service covers all five, because it only cares about one thing: did a successful ping arrive on time?
Step 1: Create a check in FlatLyne
Sign in, open Checks and select Create check. Set:
- Schedule type: Cron expression, and paste the same expression as your crontab line, for example
0 2 * * *(02:00 every day). Pick the Timezone your cron daemon uses. - Grace period: how long FlatLyne waits after the due time before marking the check Down.
For a cron check, the due time is the scheduled start time, not the end of the run. A job that pings when it finishes sits in Grace for its whole run, every time, without alerting. So make the grace period longer than your job's longest normal run. A backup that takes 5 to 25 minutes is fine with a 1 hour grace period.
See the quickstart for the full walkthrough and schedules for how FlatLyne computes due times. The Free plan includes 10 checks with a shortest schedule of 1 minute.
Copy the check's ping URL from the Ping endpoint field. It looks like this:
https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NITreat it like a password. Anyone with the URL can mark your check up or down.
Step 2: Add the ping to your crontab
The simplest version pings only when the job succeeds. && runs curl only if the job exits with 0:
PING_URL=https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NI
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL"The flags matter: -f fails on HTTP errors, -sS hides progress but keeps errors, -m 10 stops a stalled request after 10 seconds, --retry 5 retries timeouts and temporary server errors, and -o /dev/null discards the PONG reply so cron does not mail it to you. If the job fails, no ping arrives and FlatLyne alerts once the due time plus grace has passed.
Report failures immediately and time the run
To hear about a failure right away instead of after the grace period, send a fail ping. To measure duration, send a start ping first. A short wrapper does both and keeps the job's exit status:
#!/bin/bash
PING_URL="https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NI"
fl_ping() { curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL$1" || true; }
fl_ping /start
if /usr/local/bin/backup.sh; then
fl_ping ""
else
status=$?
fl_ping /fail
exit "$status"
fi0 2 * * * /usr/local/bin/backup-with-pings.shA /fail ping marks the check Down at once and sends an alert. A /start ping records the start without changing the check's state or due time, so a job that starts and then hangs still goes Down on schedule. The ping after a start ping shows the run's duration on the check's page.
If you prefer one line, send the exit code: /usr/local/bin/backup.sh; curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL/$?". Code 0 counts as success and any other number as a failure.
Each check accepts 3 pings per minute (extra pings get 429), and a start plus a final ping uses 2. See ping URLs and reliability.
Ping by slug with a ping key
Instead of one URL per check, a project ping key lets you ping by name: https://api.flatlyne.com/<ping-key>/<slug>. Give the check the slug backup-db-nightly, then:
PING_URL=https://api.flatlyne.com/<ping-key>/backup-db-nightly
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 5 -o /dev/null "$PING_URL"/start, /fail and /<exit-code> work the same way after the slug. Ping keys can be revoked and replaced, but a leaked key reaches every slugged check in the project. See ping keys.
Step 3: Set up alerts
FlatLyne alerts when a check goes Down and again when it recovers. By default it emails every member of the project. You can add Slack, PagerDuty and webhook channels from the dashboard, and the Free plan allows 2 alert channels. See alerts.
Common pitfalls
- Relative paths and missing env vars. Use absolute paths in crontab lines and scripts, and set anything the job needs in the script itself.
A && B || Con one line. IfAsucceeds but the success pingBfails on the network,Cruns and sends a false fail ping. Use theif/elsewrapper above.- Unescaped
%in crontab. Cron turns%into a newline. Escape it as\%or move the command into a script. - Grace period shorter than the run. The check sits in Grace, and a slow run turns into a false Down.
- Overlapping runs. FlatLyne does not give runs an ID, so overlapping runs mix in the ping history. Prevent them with
flock -n /tmp/backup.lock /usr/local/bin/backup.sh. - DST. Run the server and the check in UTC, or use a cron check with the same expression and timezone, and avoid 01:00 to 02:59 local time.
FAQ
What is the difference between a start ping and a success ping?
A success ping says the job finished fine and sets the next due time. A start ping only records that a run began, and leaves the status and due time alone. Use both to see how long each run takes.
What happens if the server is down and cron never runs?
No ping arrives, so the due time and the grace period pass and the check goes Down. That is the main reason to monitor from outside the machine.
Will a failed ping request break my job?
Not if you use the wrapper above. || true ignores the ping failure, and a missing ping only causes an alert, never a silent miss.
Know when your jobs stop running.
Try FlatLyne free