How to monitor a GitHub Actions scheduled workflow
A GitHub Actions schedule can stop running with no error. Ping FlatLyne from the workflow and get an alert when a run is missed or fails.
To monitor a GitHub Actions scheduled workflow, add a dead man's switch: create a FlatLyne check with the same cron schedule, store its ping URL as a repository secret, and have the workflow call that URL with curl on success and on failure. If the workflow fails, FlatLyne alerts you immediately. If it never runs at all, FlatLyne alerts you when the ping is late past its grace period.
GitHub can't tell you about the second case: a run that never starts produces no failed job.
Why scheduled workflows fail silently
A schedule: trigger has several ways to stop working without any error:
- Public repositories are disabled after inactivity. GitHub automatically disables scheduled workflows in a public repository after 60 days without repository activity. You get a notice, but it's easy to miss, and nothing fails.
- Runs can be delayed or dropped. GitHub runs scheduled workflows on a best-effort basis. During periods of high load, especially around the top of the hour, a run can start late, and GitHub's documentation warns that queued runs may be dropped.
- Cron is in UTC. The expression
0 2 * * *means 02:00 UTC, not your local time. You can't set a timezone on the trigger. - Only the default branch runs. The
schedule:trigger runs the workflow file on the repository's default branch. A schedule you add on a feature branch never fires.
A check that waits for a ping catches all of these, because it doesn't depend on the workflow being alive.
Step 1: create a check
In FlatLyne, create a check that matches the workflow's schedule. The quickstart walks through the form.
- Schedule type: Cron expression, with the same expression as the workflow, for example
17 2 * * *. - Timezone:
UTC, to match GitHub. - Grace period: longer than the workflow's longest normal run plus GitHub's start delay. For a cron check, FlatLyne expects the ping at the scheduled time, so a workflow that pings at the end of its run sits in Grace for the whole run. See Schedules. A 10 minute nightly job is safe with a 1 hour grace period.
GitHub can start scheduled runs late, so don't set a tight grace period: too short gives false alerts on busy days, too long only delays the alert.
Copy the ping URL from the check's page. It looks like https://api.flatlyne.com/<token>.
Step 2: add the pings to the workflow
Save the ping URL as a repository secret named FLATLYNE_PING_URL under Settings > Secrets and variables > Actions. Anyone with the URL can mark the check up or down, so don't commit it to the repository.
name: nightly
on:
schedule:
- cron: "17 2 * * *"
workflow_dispatch:
jobs:
nightly:
runs-on: ubuntu-latest
timeout-minutes: 30
env:
URL: ${{ secrets.FLATLYNE_PING_URL }}
steps:
- name: Ping start
run: curl -fsS -m 10 --retry 5 "$URL/start" || true
- uses: actions/checkout@v4
- name: Run the job
run: ./scripts/nightly.sh
- name: Ping success
run: curl -fsS -m 10 --retry 5 "$URL"
- name: Ping failure
if: failure()
run: curl -fsS -m 10 --retry 5 "$URL/fail" || trueHow it works:
on: scheduleruns the job on the cron.workflow_dispatchadds a Run workflow button, so you can test the whole chain without waiting for the schedule.- The start ping is recorded as STARTED and doesn't change the check's state or due time. A run that starts and then hangs still goes Grace and then Down on schedule.
- The success ping is the last step. It only runs if every earlier step passed, so a failed job never reports success.
if: failure()runs the fail ping only when an earlier step failed. FlatLyne moves the check to Down right away and alerts, without waiting for the grace period.-fsS -m 10 --retry 5sets a 10 second timeout and retries transient errors.|| trueon the start and fail pings keeps a network problem from failing the workflow. See Reliability tips.timeout-minutesmakes a hung job fail, and so send the fail ping.
Each check accepts 3 pings per minute. A run uses 2 (start and final); extra pings get a 429 and aren't recorded.
Alternative: ping keys and slugs
A project ping key lets workflows ping by name: give the check a slug such as nightly, store the key as a secret, and use https://api.flatlyne.com/<ping-key>/nightly, with /start and /fail appended as above. A ping key reaches every check with a slug in the project, so a leak reaches further than one check's URL. See Ping keys.
Step 3: set up alerts
Open Integrations in your project to choose where alerts go. Email is set up for each project; Slack, PagerDuty, WhatsApp, phone calls and webhooks are also available. FlatLyne sends one alert when the check goes down and one when it recovers, not one per missed run. See Alert channels.
To test, use Run workflow, confirm the check turns Up, then make the script exit 1 and confirm the down alert.
Pitfalls
- Secrets aren't available to pull requests from forks. If you also trigger this workflow on
pull_request,secrets.FLATLYNE_PING_URLis empty for fork PRs andcurlfails. Keep the pings in the scheduled workflow, or guard them with anif:condition on the event name. - Timezone mismatch. The workflow cron is UTC. If the check is in another timezone, the due time is wrong. Use
UTCon both sides. - Shortest intervals differ. GitHub's shortest schedule interval is 5 minutes. FlatLyne accepts 1 minute on Free and 30 seconds on Pro, so the limit is GitHub's. With a 5 minute schedule, the grace period must be larger than GitHub's start delay, or you'll get false alerts.
FAQ
Does GitHub notify me when a scheduled workflow stops running?
Not reliably when a run never starts. A failed run shows up in GitHub, but a run that is delayed, dropped or disabled leaves no failed job to notify you about. A check that expects a ping covers that gap.
How long should the grace period be?
Longer than the longest normal run, plus some slack for GitHub starting the run late. For a nightly job that takes a few minutes, 30 minutes to 1 hour works. See Schedules.
Do I need the start ping?
No. Success and fail pings detect missed and failed runs. The start ping adds run durations to the history and shows runs that hung.
Can I monitor several workflows with one secret?
Not with plain ping URLs: each check has its own, so use one secret per workflow. A project ping key covers every check with a slug.
Know when your jobs stop running.
Try FlatLyne free