How to monitor a Kubernetes CronJob

A Kubernetes CronJob can skip runs or fail without anyone noticing. Add a FlatLyne check and get alerted when it doesn't run or exits non-zero.

To monitor a Kubernetes CronJob, have the job's container ping an external service when it finishes, and alert when the ping doesn't arrive. With FlatLyne that means three things: create a check whose schedule matches the CronJob, store its ping URL in a Secret, and make the container run the job and then curl the URL with the job's exit code. If the job fails, FlatLyne marks the check Down right away. If the CronJob never runs at all, the ping goes missing and FlatLyne alerts once the grace period passes.

Why CronJobs fail silently

Kubernetes shows you the Jobs it created. It doesn't tell you about runs that never happened. Common ways a CronJob goes quiet:

  • Missed schedules. If the controller is down, or a run can't start within .spec.startingDeadlineSeconds of its scheduled time, Kubernetes counts that run as missed and skips it. No Job is created, so there is nothing to inspect afterwards.
  • Too many missed starts. If the controller counts more than 100 missed start times, it stops scheduling the CronJob and logs an error. Setting startingDeadlineSeconds limits the window in which missed runs are counted.
  • concurrencyPolicy. The default, Allow, lets runs overlap. Forbid skips a new run while the previous one is still going, and Replace cancels the running Job in favor of the new one. With Forbid, a hung job quietly blocks every run after it.
  • History limits. Finished Jobs are garbage-collected according to successfulJobsHistoryLimit and failedJobsHistoryLimit. By the time you look, the failed Job and its pod logs may be gone.
  • Timezones. Unless .spec.timeZone is set, the schedule is interpreted in the time zone of the controller manager, which is often not the one you had in mind.

An outside monitor doesn't depend on any of this. It only asks one question: did a successful run report in on time?

Step 1: create a FlatLyne check

Create a check with the same schedule as the CronJob. Choose Cron expression, enter the CronJob's schedule value, and pick the same timezone you use in .spec.timeZone (or the controller manager's). Set the Grace period longer than the job's longest normal run, because a cron check's due time is the scheduled start time, so a job that pings at the end sits in Grace for the whole run.

Copy the ping URL from the check's page. The quickstart walks through this, and Schedules, grace and timezones explains how due times are computed.

Store the URL in a Secret:

Terminal
kubectl create secret generic flatlyne \
  --from-literal=backup-db-nightly-url='https://api.flatlyne.com/CSps9LOZGccsl2o7ieL0_YrQyZJtkGK_0H1u30FJ-NI'

Step 2: ping from the CronJob

This CronJob runs one container. The shell sends a start ping, runs the job, saves its exit code, pings with that code, and exits with the job's own status so Kubernetes still sees the real result.

backup-db-nightly.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: backup-db-nightly
spec:
  schedule: "0 2 * * *"
  timeZone: "Etc/UTC"
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 600
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 3
  jobTemplate:
    spec:
      backoffLimit: 0
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: backup
              image: registry.example.com/backup-db:1.4
              command: ["/bin/sh", "-c"]
              args:
                - |
                  curl -fsS -m 10 --retry 5 -o /dev/null "$FLATLYNE_PING_URL/start" || true
                  /usr/local/bin/backup.sh
                  status=$?
                  curl -fsS -m 10 --retry 5 -o /dev/null "$FLATLYNE_PING_URL/$status" || true
                  exit $status
              env:
                - name: FLATLYNE_PING_URL
                  valueFrom:
                    secretKeyRef:
                      name: flatlyne
                      key: backup-db-nightly-url

The image must contain curl and /bin/sh. The || true after each ping means a network problem never fails the job itself. If a ping doesn't arrive, FlatLyne notices the missing ping and alerts you, so the worst case is a false alarm, never a silent miss.

-m 10 caps each attempt at 10 seconds and --retry 5 retries transient errors. See Reliability tips.

How the exit code is reported

Appending /$status to the URL is the exit-code form, /<token>/<exit-code>. FlatLyne treats 0 as a success ping and any other whole number as a fail ping, and stores the code with the ping. A fail ping moves the check to Down immediately, without waiting for the grace period. The /start ping doesn't change the check's state or due time: it only records when a run began, so a run that starts and then hangs still goes late and down on schedule.

If you'd rather choose the URL yourself, use /fail explicitly:

Explicit success or fail
if /usr/local/bin/backup.sh; then
  curl -fsS -m 10 --retry 5 -o /dev/null "$FLATLYNE_PING_URL" || true
else
  status=$?
  curl -fsS -m 10 --retry 5 -o /dev/null "$FLATLYNE_PING_URL/fail" || true
  exit $status
fi

More patterns, including sending the job's output as the ping body, are in Failures, exit codes and start pings.

Alternative: ping keys and slugs

Instead of one Secret per check, a project ping key lets every CronJob ping by name: https://api.flatlyne.com/<ping-key>/<slug>, with the same /start, /fail and /<exit-code> suffixes. Store the key in one Secret, set a slug on each check, and build the URL from both. Ping keys can be revoked and rotated, and on Pro a check's own ping URL can be regenerated too. See Ping keys.

Step 3: set up alerts

By default FlatLyne emails every member of the project when a check goes Down and when it recovers. To send alerts to Slack, PagerDuty or a webhook as well, add them under Integrations. One outage sends one alert, however many runs fail in a row. See Alert channels.

To test, run kubectl create job --from=cronjob/backup-db-nightly test-run and confirm the check turns Up.

Pitfalls

  • The image has no curl. Minimal and distroless images often don't. The job then runs, the ping fails silently because of || true, and the check goes late. Use an image that includes curl, or build one that does.
  • DNS or egress is blocked. Pings need outbound HTTPS to api.flatlyne.com on port 443, and FlatLyne doesn't publish IP ranges. If a NetworkPolicy or egress proxy restricts traffic, allow the hostname. If you use a proxy, curl reads the https_proxy variable.
  • The pod is killed before the final ping. Node drains, evictions, or an activeDeadlineSeconds timeout can end the pod before it reaches the last curl. That's the case the missed-ping alert exists for: the check goes Down once the grace period runs out.
  • Retries cause duplicate pings. With restartPolicy: Never, a non-zero backoffLimit creates a new pod for each failed attempt, and each attempt sends its own pings. The first failure marks the check Down and alerts, and a later success sends a recovery. Set backoffLimit: 0 if you want one run, one result.
  • The rate limit. Each check accepts 3 pings per minute. Extra pings get 429 and aren't recorded. A run that sends a start ping and a final ping uses 2, so fast retries can hit the limit.

FAQ

Does a failed CronJob pod trigger an alert?

Only if it reports the failure. If the container exits non-zero after your wrapper has sent /$status, the check goes Down immediately. If it is killed before the final ping, FlatLyne alerts when the grace period passes.

What happens if the CronJob never creates a Job?

No ping arrives, so the check goes Grace and then Down after the due time plus the grace period. This covers skipped runs, a suspended CronJob, and a controller that stopped scheduling.

Do I need a sidecar?

No. A single container with sh -c is enough as long as the image has curl.

Can one check cover several CronJobs?

Give each CronJob its own check. A check tracks one schedule, so a shared check can't tell you which job stopped.

Know when your jobs stop running.

Try FlatLyne free