The Watchbillby LatticeDDI

How to reduce false uptime alerts

Published . Updated .

This guide explains how to stop paging a technician for a blip, a challenge page, or a single dropped request, without hiding a real outage until morning.

Require an outage to be confirmed before alerting, either by consecutive failed checks or by agreement across multiple check locations, alert only on state changes, and route by severity. Each step trades a little detection time for far fewer 3 a.m. false alarms.

A false positive here means the alert said down and the site was usable, or the failure was not the failure you meant to watch. People stop reading alerts that lie. The dangerous outcome is not the extra page. It is the real page that gets snoozed because last night's was nothing.

Why does a single check lie?

Networks drop packets. A resolver has a bad moment. A CDN asks a strange client to solve a challenge and returns a page that is not the site. A deploy restarts a process for ten seconds between two of your checks, or during one of them. Any of those can fail one request and succeed the next.

RFC 9110 separates status codes into classes. A 200 and a 500 are different facts. A 401 or a 403 is the server refusing the client, which may be the site protecting itself from your checker. A 429 is the server asking for less traffic. Treating every non-200 as down will page you for a policy you could have read. The status code note goes through the classes. This page is about when to speak.

One failed check is a sample. It is not yet an incident.

What is the difference between consecutive failures and a location quorum?

Consecutive failures means you wait until the check has failed on purpose, more than once, with no success in between. Two failures a minute apart means the problem lasted about a minute, which is longer than a single dropped packet. Three failures means you waited longer and you will be more sure. The cost is the wait. If you require two minutes, you will not open an incident in the first minute.

A location quorum means several checkers in different places have to agree. One location fails and the others succeed: you stay quiet, because the fault may be the path from that one place. Two of three fail: you alert. That catches a regional network problem that a single location would call an outage, and it misses a failure that only one region can see. It also costs you more checkers.

Both are confirmation. They answer different doubts. Consecutive checks doubt the clock. Multiple locations doubt the path. You can use either, or both. Say which one you use in the client agreement, because the minute you waited is part of the downtime math in the calculator if your contract starts at the first failure instead of at the confirmed one.

MethodWhat it doubtsWhat you give up
One failed checkNothing. It alerts immediately.Sleep, and trust in the next alert.
Two or more consecutive failuresA one-off blip on a single path.The time between the first failure and the confirming one.
Agreement across locationsA fault that only one checker can see.Outages visible from only one region, plus the cost of more checkers.

When should an alert fire?

Fire when the state changes, and do not fire again until it changes back. Down, then down, then down, is one incident. A message every minute retrains the on-call to mute the sender. A recovery message matters as much as the down message. It closes the loop and, if your tickets follow the subject, it is what lets a rule resolve the ticket. That path is in alerts into tickets.

Route by severity only if you have a real difference in severity. A brochure site and a payment site should not wake the same person the same way. If every check is "critical," the word does nothing. Two routes are enough for a small shop: a mailbox that becomes a ticket during the day, and a page for the few names that lose money while they are down.

Measure the fatigue. Count alerts per week, how many were real, and how many were followed by a recovery inside five minutes. If the short ones dominate, lengthen the confirmation. If real outages are arriving late, shorten it, and accept more noise on purpose. Write the count down. A feeling that "alerts are bad" is not a tuning decision.

What should you exclude on purpose?

A scheduled maintenance window is a time you already told the client about. Alerting during it is a false alarm you scheduled. How to write that exclusion is the maintenance window note. Do not silence a check forever because it was noisy once. Silence it for the window, then turn the rule off.

A challenge page, a login wall, and a rate limit are measurements you refused, or should refuse. Decide that before the first 3 a.m. Otherwise the on-call will decide it while half awake, and the decision will be "mute."

How The Watchbill helps

The Watchbill confirms outages with time rather than a vote across locations. A site is marked down or up only after the same result on two consecutive minutes, so one dropped request never opens an outage. A result more than three minutes old isn't treated as the site's current state, so a stalled check can't keep showing an old answer.

Alerts go out when the confirmed state changes, once for the outage and once for the recovery, rather than every minute. A bot-challenge page, a 401, a 403, or a 429 is recorded as a measurement the site refused, not as an outage, so a firewall rule doesn't page anyone at 3 a.m. Checks run on Cloudflare's network, placed by region on a best-effort basis, and results aren't labeled with a city. If you need a quorum across named locations, that's a different kind of tool, and the table above explains the trade-off. To see how quiet the alerts are on your own sites, create a free account and add a few checks.

Sources

  1. RFC 9110, HTTP Semantics, status codes. Accessed October 10, 2026.

Related guides