If you run a dedicated or Robot server on Hetzner, you have almost certainly seen SysMon: it is free, it is one form in the Robot panel, and it is the first thing most people set up because it is right there. It is also not the whole answer, and the gap between "SysMon is green" and "my server is actually fine" is exactly where outages hide.
What SysMon actually checks
SysMon runs from Hetzner's own monitoring pool and supports a fixed set of protocol checks against your server: ping, TCP/UDP port, HTTP (accepting only 200, 301, or 302 as healthy), FTP, POP, IMAP, SMTP, and DNS. You set the interval and a reminder cadence, and it requires two consecutive failures before it alerts, which cuts false alarms at the cost of a delay on the first real one.
That is a genuinely useful floor. A port check on 443 catches "nginx crashed and nothing is listening." An http check catches a web server that answers but returns a 5xx. For a single box with no redundancy, having any external eyes on it beats having none, and SysMon costs nothing to turn on.
The two things it structurally cannot tell you
It only proves a status code, not that the app works. SysMon's HTTP check accepts 200, 301, or 302 and nothing else. A page that returns 200 with a stack trace instead of your homepage, a health endpoint that always returns 200 regardless of what is actually broken behind it, or an API that answers 200 with an empty or malformed body all read as healthy to SysMon. It checks that something answered with an acceptable code, which is a real signal and a narrow one.
It checks from one place: Hetzner's own network. SysMon's monitoring pool queries your server from Hetzner's infrastructure (its own docs describe granting pool.sysmon.hetzner.com access through your firewall). That is the same provider, often the same data center region, as the server being checked. A routing problem between your box and the rest of the internet, a firewall rule that blocks traffic from outside Hetzner's own ranges but not from inside them, or a regional peering issue that makes your site slow or unreachable for customers in Singapore while it answers Hetzner's own probe in Nuremberg instantly: none of these show up, because SysMon never leaves the neighborhood.
Both gaps point at the same underlying fact: SysMon confirms the server is up from the vendor's point of view. It cannot confirm your customers, who are not on Hetzner's network and are not satisfied by a bare 200, are actually getting a working site.
Setting up an outside check
The fix is not to replace SysMon; keep it, it is free and it catches the box-is-dead case fast. Add a check that runs from outside Hetzner entirely, against what a real visitor sees, on top of it.
An HTTP check against your real endpoint, from several regions. On RealUptime Monitor, a check hits the actual URL a visitor loads, not just any 200-yielding path, and you can assert on more than the status code: response time thresholds, and a specific string or pattern in the body if you want to catch "the page loaded but the content is wrong." Running from multiple regions turns a single green/red into a map: if a check fails from Singapore but passes from Virginia and Frankfurt, that is a routing or peering problem specific to one path to your server, not a dead box, and you know which one to chase.
A TCP check for anything that is not HTTP. If the server also runs Postgres, Redis, an SMTP relay, or an SSH bastion on a non-standard port, a TCP check confirms the port accepts connections from outside your own network, which is a different and complementary question to whether nginx is up.
A status page, if anyone besides you needs to know. Hetzner's Robot panel has no way to tell your team or your customers what is happening. A public or internal status page turns "I noticed the alert" into "everyone affected already knows," without a Slack message you have to type by hand mid-incident.
SysMon's alert delay, and why it exists
SysMon's two-consecutive-failure rule is worth understanding on its own terms rather than treating as a flaw. A single failed check is cheap: a momentary network blip on Hetzner's monitoring pool side, a packet lost between their probe and your box, a check that happened to run during a brief CPU spike from a cron job. If every failed check paged you, most pages would be noise, and noise is how real alerts get ignored. Requiring the second consecutive failure before Hetzner notifies you filters that noise out, at the cost of however long your check interval is, twice over, before you hear about a real outage.
An outside HTTP or TCP check follows the same logic for the same reason, and it is worth setting the interval and the failure threshold deliberately rather than leaving the defaults unexamined. A one-minute interval with a two-failure threshold means at most two minutes between a real outage starting and an alert firing; a five-minute interval with the same threshold means up to ten. For a server that only you depend on, five minutes might be fine. For one that customers hit directly, the tighter interval is worth the marginally higher check volume.
A troubleshooting order that actually saves time
When an alert fires, whether from SysMon or from an outside check, the order you check things in matters more than how fast you check them. A reasonable sequence:
- Check whether SysMon and the outside check agree. If both are down, the problem is almost certainly the box itself: the process crashed, the server is out of memory, or the network interface is down. SSH in (or use the Robot panel's remote console if SSH itself is unreachable) and look at
systemctl status,docker ps, or whatever supervises your app. - If only the outside check is down and SysMon is green, suspect the network path, not the process. A firewall rule that changed, a DNS record pointed somewhere stale, or a regional routing issue between your outside monitor and the server are the likely causes, since the process itself is answering Hetzner's own probe fine.
- If only SysMon is down and the outside check is green, suspect the reverse: something specific to Hetzner's monitoring pool reaching your box, most often a firewall rule that was tightened without adding the SysMon IP ranges back in.
This is the same reasoning covered in more general form in what a TCP port check proves, and what it does not: a check that only tells you "up" or "down" from one vantage point cannot distinguish a dead process from a broken path to it. Two checks from two different vantage points can, and the disagreement between them is itself the diagnosis.
What this looks like together
A reasonable setup for a single Hetzner box running a web app and a database on the side: keep SysMon's port check on 443 and 5432 as the fast, free, close-to-the-metal layer that catches "the process died" in under a minute. Add an HTTP check on RealUptime against the real app URL from at least two regions outside Hetzner's network, so you catch routing and content problems SysMon's status-code check cannot see. Add a TCP check on the database port if anything outside the box connects to it directly. If the server does scheduled work, a nightly backup, a cron-driven export, add a heartbeat monitor so a silently failing cron job pages you instead of going unnoticed for a month.
None of this requires installing anything on the server for the HTTP and TCP checks; they are outbound-only from RealUptime's side, pointed at whatever is already public. If you want host-level checks too (disk, memory, or a service that only listens on localhost), the Monitor agent runs as a small process on the box itself and reports alongside the outside checks, which is the same pattern people running Coolify on Hetzner already use, covered in our Coolify monitoring guide.
The short version
SysMon is a real, free monitoring layer and worth keeping. It checks from inside Hetzner's own network and only as far as a status code; it cannot tell you whether the app behind that status code works, or whether a customer somewhere else on the internet can actually reach the box. An external HTTP or TCP check, run from regions Hetzner does not control, closes exactly that gap, and takes a few minutes to set up on top of what you already have.