Why Do Status Pages Say Operational During an Outage?

Farouk Ben. - Founder at OdownFarouk Ben.()
Why Do Status Pages Say Operational During an Outage? - Odown - uptime monitoring and status page

Why Do Status Pages Say Operational During an Outage? The Four Green-Light Failures

Most status pages say operational during an outage because a person has to change them and that person is currently busy fixing the outage. The remaining explanations are structural: the page is driven by checks that run inside the same infrastructure that failed, the component list is too coarse to represent what actually broke, or nobody is rewarded for turning the light red. Deliberate concealment is real but it is the least common of the four.

This article covers each of those four failures, why the manual one is so much more common than users assume, what an honest status page looks like, and how to read someone else's page when you depend on their service. It is written from the position that the incentive problem is real and that saying so is more useful than pretending otherwise.

The first failure: somebody has to press the button

Most status pages are updated by hand. An incident begins, responders start working, and the status page is a separate system requiring a separate login and a decision about wording. The people who could update it are the people least available to do so, and the update competes directly with the fix.

There is a second-order effect that makes this worse than simple busyness. Early in an incident nobody is certain what is broken, and a status update is a public statement. Posting investigating elevated errors three minutes in feels premature when it might turn out to be one region or one customer. So the responder waits for clarity, clarity takes twenty minutes, and by the time the page turns yellow your users have been staring at a green dot for a third of an hour while the site refused to load.

The fix is organisational rather than technical: someone whose job during an incident is communication and nothing else, empowered to post before the cause is known. That role exists in mature incident processes precisely because leaving it to the responders means it does not happen.

The second failure: the check runs inside the failure

Automated status pages avoid the human problem and introduce a subtler one. If the check that drives the page runs in the same datacentre, cloud region, or network as the service it watches, it can only report failures that do not affect itself.

An internal health check that queries the application over the local network will happily report success while your CDN configuration, DNS records, TLS certificate, or edge firewall makes the service unreachable to everyone outside. All of those failures are invisible from inside, and all of them are total from the customer's point of view. The status page is not lying. It is answering a different question from the one users are asking, which is not is the process running but can I use this from where I am.

The related version is scope. A check against the homepage proves the homepage renders. It says nothing about whether login works, whether the API is returning errors, or whether checkout is timing out at the payment step. The status page examples worth studying generally distinguish between components at the level users would notice, which is the point of having components at all.

The third failure: the component list is too coarse

A status page with three components, labelled Website, API, and Dashboard, cannot represent a partial failure, and partial failure is the normal shape of an outage.

If checkout is broken but browsing works, is the Website operational? A binary component forces a bad answer either way. Marking it down overstates the impact and triggers escalations for customers who are unaffected. Marking it up understates it for the people who cannot complete a purchase. Most teams pick the second, because the first generates complaints from people who are fine, and the accumulation of that choice is a page that stays green through a lot of real damage.

Degraded performance as a middle state helps but only if it is used. It tends to become the default for anything ambiguous, which over time turns yellow into a colour meaning something might be happening, and users learn to ignore it exactly as they learned to ignore green.

The structural fix is components that map to what users do rather than to how the system is built, and enough of them that a partial failure has somewhere to be reported honestly.

The fourth failure: nobody is rewarded for red

This one deserves to be stated plainly rather than implied.

Status pages carry commercial weight. They are cited in sales conversations, referenced against contractual uptime commitments, screenshotted by competitors, and read by customers deciding whether to renew. That creates a consistent pressure in one direction: a component marked down is a permanent public record, and a component left up is not.

This rarely takes the form of anyone deciding to lie. It takes the form of a very high evidentiary bar for turning something red and a very low one for calling it resolved, of incidents backdated to start later than they did, of regional outages classified as not qualifying for a status change, and of an ambiguous case resolving in the direction that requires no announcement. Each decision is individually defensible. The aggregate is a status page that under-reports.

The counter is straightforward and a genuine competitive advantage: publish uptime measured by something you do not control, define in advance what triggers a status change, and post first and refine later. Vendors who do this get noticed, because the contrast is sharp.

Common mistakes in running a status page

Leaving updates to the people fixing the incident. They are the busiest people in the building and every minute spent on wording is a minute not spent on the fix. Assign communication to someone with no repair responsibilities.

Driving the page from checks inside your own infrastructure. Anything that fails at the network edge, in DNS, or at the certificate layer is invisible to an internal check while being total for users. Measure from outside.

Using too few components to describe a partial failure. Three broad components force a binary answer to a question that is rarely binary, and the safe-looking answer is almost always the one that understates impact.

Waiting for certainty before posting. Users can already see that something is wrong. An early post saying you are investigating costs nothing and buys enormous credibility compared to twenty minutes of green.

Setting a higher bar for red than for green. If turning a component down requires proof and turning it back up requires a hunch, the page drifts optimistic regardless of anyone's intent. Define both thresholds in advance.

FAQ

Why does a status page say operational when the site is down?

Most often because updating it is a manual step and the people who could do it are working on the outage. The structural causes are checks that run inside the failing infrastructure, component lists too coarse to show a partial failure, and the commercial pressure that makes green the safer default.

Are status pages updated automatically?

Some are, many are not, and hybrids are common. Automation removes the delay but introduces its own problem: if the checks driving the page run inside the same infrastructure as the service, they cannot detect failures at the edge, in DNS, or at the certificate layer.

Can I trust a vendor's status page?

Treat it as one signal rather than the answer. It tells you what the vendor has acknowledged, which is not the same as what is happening. If a service matters to your business, monitor it yourself from outside and compare.

How should I check whether a service is actually down?

Test it from outside the vendor's network, ideally from more than one region, against the specific function you use rather than the homepage. Partial and regional failures are the common case, and both look fine from a single check against a landing page.

Closing thought

The uncomfortable part of this question is that the honest answer is not mainly about dishonesty. It is about a page maintained by hand during the exact hour nobody has a spare hand, measured by systems inside the thing that broke, using categories too blunt to describe what actually happened, in a commercial context where one colour is expensive and the other is free. Any one of those would produce optimistic reporting. Together they produce a green dot above a site that will not load, and the vendors who fix it are the ones who decided to be measured by something they do not control.

That is what an outside-in check provides. Odown runs checks from seventeen global locations at intervals down to one minute on every plan, including the twelve dollar tier, and status pages run on your own domain over HTTPS. Because the checks originate outside your infrastructure, the failures that internal health checks structurally cannot see are the ones they report, which is the difference between a status page describing your servers and one describing your users' experience.