Skip to content
Knowledge base

How to report a failed or stuck cron job run, not just a missed one

A plain heartbeat only notices a job that stops checking in. Add a start ping and a fail ping and RealUptime sees every run: how long it took, whether it failed and with what exit code, and whether it started and then hung.

1. Know what the plain ping cannot tell you

The plain ping means the job succeeded. If the job runs and fails, it simply never sends that ping, so the monitor only goes down once the interval and grace period have run out, and the alert says the job went quiet rather than that it failed. If the job starts and hangs, the same thing happens, later. Two extra endpoints close both gaps without changing anything for a monitor that never uses them.

2. Send a start ping and a fail ping

Request https://ingest.realuptime.io/api/ping/<token>/start when the job begins, the plain https://ingest.realuptime.io/api/ping/<token> when it succeeds, and https://ingest.realuptime.io/api/ping/<token>/fail when it fails. In a crontab that is: curl the start URL, run the job, then && curl the plain URL || curl the fail URL. A failed run marks the monitor down at once with "run failed" wording, and the next successful run recovers it. All three share the same 120 requests per minute per token.

3. Send the exit code with the failure

POST a body with the fail ping to say why: a bare number is recorded as the exit code, and any other text as a short message, cut to 500 characters. It shows in the monitor's run history and in the alert, so the person who gets paged knows whether it was exit code 2 or a full disk before opening a terminal.

4. Catch a run that started and never finished

On the monitor's page, set a maximum runtime: how long one run may take, from 1 second to 24 hours. A run that has sent its start but no success or fail within that time is marked stuck, and the monitor goes down once with "run stuck" wording that names the limit. The next successful run recovers it. It needs the start ping, and it is off until you set it.

5. Pair overlapping runs, and know what happens when a job keeps failing

If two runs of the same job can be in flight at once, add the same rid query parameter (up to 64 letters, digits, dots, colons, hyphens or underscores) to a run's start, success and fail requests so each finish pairs with its own start; without one, a finish closes the most recent run. A job that fails and recovers three times in 24 hours is held down on the next failure instead of alerting on every run: further successes are recorded but don't clear it, and it recovers after one full interval with no failed or stuck run. Pausing the monitor clears the hold.

Go deeper

The full reference lives in the docs: Heartbeat monitors documentation. Error codes named above are each explained in the error-code reference.