The Watchbillby LatticeDDI

Should maintenance count on an SLA?

Published . Updated .

This guide explains how to take a site down on purpose without paging your own team or using up SLA time you didn't mean to spend.

Schedule maintenance windows in advance, tell clients, suppress outage alerts only for the affected checks during the window, and decide in the contract whether scheduled downtime counts against the SLA. Most SLAs exclude announced maintenance, but only if notice rules are met.

Maintenance is work you chose. An outage is work that chose you. If you mix them in the same alert stream, the on-call cannot tell which pager events they are allowed to ignore. If you mix them in the same percent, the client will argue about a Saturday you already agreed to. NIST's writing on resilient systems treats availability as something you design for, including the times you take a component out. The contract is where you write that design down in a sentence a non-engineer can apply.

What is a window, exactly?

A window has a start, an end, a list of checks, and a reason a person can repeat. "Saturday night, the booking site, database upgrade" is a window. "Sometime this weekend, if we get to it" is not. The end matters because suppression that never expires is how a real outage the next morning stays silent.

Write the window in the client's timezone and say the zone. A start of "midnight" without a zone is a bet. Put a buffer at the end. If the work takes two hours, the window is three, and you close it early if you finish early. An open-ended suppression is not a buffer.

If several checks share one site and you are only touching the API, suppress the API check. Leave the homepage check live. A window that silences every URL the client owns will hide a failure you did not cause.

How much notice is enough?

The notice is whatever the contract says. If the contract is silent, you do not have an exclusion yet. You have a hope. A common shape is one full business day for ordinary work and a shorter notice for an emergency change you will explain after. Pick numbers you can keep. A promise of ten days that you break every month trains the client to ignore the notice, and then the exclusion is the first thing they dispute.

Send the notice in the channel the client already reads, and put the same sentences on the status page. The page is what they will open when something looks down. Wording for that page is in the status page note. Say what will be unavailable, when it starts, when it should end, and whether they need to do anything. Most of the time they need to do nothing, and you should say that.

What may the alerts ignore?

Ignore the checks named in the window, between the start and the end, and only those. A down alert for a different client, or for a different URL, still means what it usually means. Recovery alerts can stay on. A recovery during the window tells you the work finished, which is useful, and it should not open a second incident.

Do not suppress by muting the whole account. The next person on call will not know the mute is there. A window tied to a check and a clock is something you can read on Monday.

When the window ends, the next failed check is a real alert again. Confirm that on purpose. Run the check, or watch the first result after the end, before you go to bed. A window that ends into a still-down site should page someone. That is the point of the end time.

DecisionA clear ruleA rule you will regret
Which checksOnly the URLs you are changingEvery check on the account
How longStart and end, with a zoneUntil someone remembers to unmute
What still alertsAnything outside that listNothing, because the mute was account-wide
After the endThe next failure pages as usualThe mute is still on at breakfast

Does the window count against the SLA?

Decide in the agreement, in the same sentence as the percent. A workable exclusion says that scheduled downtime is outside the percent when you gave the notice the agreement requires, you stayed inside the window, and you named the service. Miss the notice, run long, or take down a service you did not name, and those minutes count.

The arithmetic is in the downtime calculator. A 99.9% month of 30 days allows 43 minutes 12 seconds. A two-hour upgrade consumes that budget several times over if it counts, and it consumes none of it if the exclusion holds. Write which one you sell. Do not leave it to the quarterly review.

Emergency work can have its own rule: it counts unless you were responding to a failure that would have counted anyway, or it never counts, or it counts above a cap. Any of those is a policy. Silence is how you end up debating it on a Sunday.

What do you say on the status page?

Before the window: what, when, and that you expect it. During the window: that the work is in progress, so a visitor does not file a second ticket. After the window: that it ended, and the current state. Take the maintenance line down when the window ends so the page does not say "work in progress" on a Tuesday.

If the work fails and the site is still down at the end, the page changes from a maintenance line to an incident line. Those are different sentences. Do not stretch the maintenance banner to cover a failed change. The banner was the exclusion. The incident is the outage.

The percent you report should match the exclusion you wrote. If confirmation waits for two failed minutes, remember that those minutes are part of the measurement story in false alarms. Tell the client which clock the report uses.

How The Watchbill helps

The Watchbill doesn't store maintenance windows, and it doesn't mute alerts on a schedule. During planned work, expect a Down message once two consecutive minutes fail and an Up message when the site comes back. Knowing that ahead of time is useful. If you tell the on-call tech that both messages are coming, the Up message becomes the confirmation that the work finished and the site is answering again.

The rest of this guide still applies around it. The window, the notice, and the SLA exclusion live in your contract and your messages to the client. The 90-day check history in your account shows when the site was down, so you can match those minutes against the window you announced when you write the monthly report. On Pro and Business, the Down and Up messages can also go to your PSA as ticket email, so the planned outage is recorded on a ticket you can note and close. Plans are on the pricing page.

Sources

  1. NIST SP 800-160 Vol. 2 Rev. 1, availability as a concern of a resilient system. Accessed October 10, 2026.

Related guides