Uptime monitoring glossary for MSPs
Published . Updated .
This glossary defines the words that show up in an uptime report, a client agreement, and a ticket about a site that stopped answering. Each entry is the meaning, then the detail you need before you use the word with a client.
Uptime monitoring is automated checking, at a fixed interval, that a website or service answers correctly from outside your network, with an alert when it stops. It measures availability from the visitor's side. It does not measure server health itself.
Use the words the same way in the proposal and in the monthly report. A client who hears "SLA" in the sale and "uptime" in the report will assume they are the same number. Often they should not be. The table is the split to keep straight. The longer arithmetic is in the downtime calculator.
| Word | Who it binds | What it is allowed to mean |
|---|---|---|
| SLI | Your report | The measurement you can show |
| SLO | Your team | The internal target |
| SLA | The contract | The promise that has a remedy |
Before you send a report, check the words in it.
- The percent names a period and a method.
- The status page uses the client's name for the service.
- The alert says down only after your confirmation rule.
- The ticket subject matches the rule your PSA parses.
- A term you cannot define in one breath gets cut from the client email.
Uptime
Uptime is the share of a chosen period when the service counts as working. Write it as a percent of that period, such as a calendar month or 365 days, and say whether you counted recorded minutes. It is not a feeling about the server, and it is not CPU load. Minutes with no check are a gap. Call them a gap, or the percent will look calmer than the client's week felt.
Availability
Availability is uptime expressed as a percent of the period you agreed to measure. People use the two words as synonyms. In a contract, availability is the figure with a denominator: which minutes, which URL, and which failures count. A higher percent is a smaller allowance for failure. The allowance is the period length times one minus the percent, which is what the downtime calculator works out.
SLA
An SLA, a service level agreement, is the availability promise written into a contract, usually with a remedy such as a service credit. It is not the number you hope for. It is the number you are willing to attach money to. Notice rules, a cap on the credit, and an exclusion for announced maintenance belong in the same clause. A percent with none of those is a slogan with a digit in it.
SLO
An SLO, a service level objective, is the internal target your team uses to decide whether the service is healthy enough. It can be tighter than the SLA so you notice trouble before the contract does. It does not, by itself, owe the client a credit. When the SLO and the SLA use different periods, label them. Otherwise the quarterly review becomes an argument about a spreadsheet.
SLI
An SLI, a service level indicator, is the measurement behind the percent. For an uptime check it is usually the share of recorded checks, or recorded minutes, that met your definition of up. The indicator is only as honest as the check. A one-minute HTTPS request is an indicator for that URL on that timer. It is not an indicator for every page, every region, or every second.
Check interval
The check interval is how long you wait between probes. A one-minute interval cannot see an outage that starts and ends between two probes. A five-minute interval misses more and costs less. Shorter is not free: you pay in load, in noise, and sometimes in the vendor's price. State the interval next to any percent you publish, because the interval is part of what the percent means.
Synthetic monitoring
Synthetic monitoring is a scripted probe you send on a schedule, as opposed to waiting for a real visitor's browser to report an error. An HTTPS check of a homepage is the small end of synthetic monitoring. A multi-step transaction that logs in and buys a ticket is the large end. Both are outside-in. Neither replaces the server's own metrics. They answer "could a visitor complete this," on the timer you set.
Outage confirmation
Outage confirmation means you wait for more than one failed sample before you call an incident. Consecutive failures wait on the clock. A quorum waits until several locations agree. Both trade a little detection time for fewer false pages. Say which one you use, because the minutes you waited are either inside the downtime figure or they are not. One lonely failure is a sample, not yet an outage.
False positive
A false positive is an alert that says the service is down when the failure was a blip, a blocked checker, or a challenge page the visitor never saw. The cost is trust. After enough false pages, a real page gets snoozed. You reduce them with confirmation, by alerting only when the state changes, and by refusing to call a 401, a 403, or a 429 an outage until you know visitors are blocked too.
Status page
A status page is the notice you give a client about current state, the start of an open incident, recent history, and how to reach you. For an MSP it should be one link per client, kept private until you publish it, and written in the client's words. It is not a dashboard of your other customers. Update it before you return the call, or the page will contradict you.
MTTR
MTTR, mean time to restore, is the average time from the start of an incident to the time the service counts as up again. Define both ends. If your clock starts at the confirmed alert and stops at the recovery message, say so. If the client starts the clock at the first failed minute, your average will look shorter than theirs. Use the same ends in the report that you used in the agreement.
MTTD
MTTD, mean time to detect, is the average time from the start of a failure to the time you know about it. On a one-minute check with two-minute confirmation, detection cannot be faster than that design. Do not publish an MTTD of "seconds" if your probe runs once a minute. The figure is a property of the interval and the confirmation rule, not of how quickly a person reads the mail.
TLS certificate
A TLS certificate is the dated credential a browser uses to decide that the name in the address bar matches the server it reached, and that a trusted authority signed it. When it expires or the chain is wrong, the browser warns. The site can still be running for someone who clicks through. The CA/Browser Forum schedule includes a 47-day maximum for certificates issued on or after 15 March 2029.
ACME
ACME is the protocol in RFC 8555 that a client uses to ask a certificate authority for a certificate and to prove control of the name. When it works, renewal is a job. It stalls on appliances, on panels you cannot automate, and wherever the HTTP or DNS challenge cannot see the name. A failed ACME job leaves the old certificate in place until it expires, so the success you need is a new date on the public hostname.
Domain expiry
Domain expiry is the end of the registration period for a name. The website and the mail records are data under that name. When the registration lapses, DNS is interrupted and both stop. Renewing the hosting invoice does not renew the name. ICANN's Expired Registration Recovery Policy requires reminder mail about one month and one week before expiry for generic top-level domains, sent to the registrant contact. A stale contact is a missed reminder.
Redemption grace period
The redemption grace period is the 30 days most unsponsored gTLD registries must offer immediately after a registrar deletes a registration, under ICANN's Expired Registration Recovery Policy. During those 30 days DNS stays off, transfers are prohibited, and the former holder can restore the name through the same registrar, usually for a fee above a normal renewal. Sponsored gTLDs are exempt. Country-code domains follow their own registries. After the period, someone else can register the name.
DNS TTL
DNS TTL, the time to live, is how many seconds a cache may reuse an answer before asking again. RFC 1034 and RFC 2181 describe that timer. A short TTL lets a change spread quickly, including a bad edit. A long TTL means a fix you publish now stays invisible to some resolvers until the old answer ages out.
NS, MX, and TXT records
NS records name the servers allowed to answer for the domain. A surprise change is a change of custody. MX records name the servers that accept mail for the domain. TXT records carry text, including SPF and DMARC, which say who may send mail as that name. Watch those three before you watch every host record in the zone. A homepage check will not notice an MX change if the website address stayed put.
Maintenance window
A maintenance window is a start and an end, in a named timezone, for named checks, when you take the service down on purpose. Alerts should go quiet only for those checks and only inside that clock. Whether the minutes count against the SLA is a contract sentence, and it holds only if you gave the notice the contract requires. A mute with no end is not a window. It is how Monday's outage stays silent.
Webhook
A webhook is an HTTP request your monitor sends to an address you control when something happens, usually with a JSON body. You can verify it and then open a ticket through your PSA's API. Retry when the receiver is down, and stop after repeated failures. Email-to-ticket is the simpler path. A webhook is the step you take when subject-line rules keep missing.
HMAC signature
An HMAC is a keyed hash, defined in RFC 2104, so a receiver can check that a body came from someone who shares a secret and that nobody edited the body. A common webhook pattern is HMAC-SHA256 over a timestamp, a dot, and the raw body. Recompute the hash and reject a mismatch. The secret is a password. Rotate it when someone who knew it leaves.
Email-to-ticket
Email-to-ticket is the PSA feature that turns mail to a given mailbox into a ticket, with the company and priority taken from rules. A stable subject, state first and check name after, lets the recovery message find the open ticket instead of opening a second one. It is the fast path into ConnectWise, Autotask, HaloPSA, or Syncro. Confirm the menu name in that vendor's docs before you train on it.
PSA
A PSA, a professional services automation system, is where an MSP keeps tickets, time, and the client list. ConnectWise, Autotask, HaloPSA, and Syncro are the names most shops mean. The monitor does not have to log into the PSA if the PSA already makes tickets from mail. Techs live in the PSA all day. An alert that never becomes a ticket is an alert that waits on someone remembering a second inbox.
RMM
An RMM, a remote monitoring and management tool, is the agent on the client's PCs and servers. It sees CPU, disks, and services inside the network. Uptime monitoring sees the public URL from outside. An RMM can say the server is up while DNS points elsewhere. An uptime check can say the URL is down while the server itself is healthy. Use both when the client can fail in both ways.
Origin and edge
The edge is the proxy or CDN that answers the public hostname. The origin is the server behind it. A normal HTTPS check of a proxied name often measures the edge, which can be green while the origin is sick. If the agreement says you watch the origin, the check has to prove the origin responded. If you watch the public URL, say that on the status page.
How The Watchbill helps
The Watchbill puts several of these definitions to work. Each check is an HTTPS request to one URL about every 60 seconds, and a site is marked down or up only after two consecutive minutes agree, which is the outage confirmation described above. Email alerts are included on every plan. Each check also records whether the site's certificate was trusted and when its domain expires, and each client can have a status page that stays private until you publish it. On Pro and Business, alerts can also go out as ticket email, a signed webhook, a Microsoft Teams message, or a PagerDuty alert. Plans are on the pricing page, and you can try it on your own sites with a free account.
Sources
- RFC 9110, HTTP Semantics. Accessed October 10, 2026.
- ICANN Expired Registration Recovery Policy (21 February 2024). Accessed October 10, 2026.
- CA/Browser Forum Baseline Requirements, section 6.3.2. Accessed October 10, 2026.
- RFC 8555, Automatic Certificate Management Environment (ACME). Accessed October 10, 2026.
- RFC 2104, HMAC. Accessed October 10, 2026.
Related guides
- How much downtime is 99.9% uptime?
- How to keep SSL certificates current
- What happens when a domain expires
- How to hear when DNS records change
- What a client status page should show
- How to reduce false uptime alerts
- Should maintenance count on an SLA?
- How uptime alerts become PSA tickets
- Pingdom alternatives for MSPs
- How an MSP picks an uptime monitor
- Which HTTP codes count as downtime