What Is Alert Deduplication and Why It Matters

Farouk Ben. - Founder at OdownFarouk Ben.()
What Is Alert Deduplication and Why It Matters - Odown - uptime monitoring and status page

What Is Alert Deduplication and Why It Matters

Alert deduplication is the mechanism that collapses many notifications about the same ongoing problem into a single alert. A monitoring system checking every minute will detect a failed service sixty times an hour, and without deduplication that is sixty pages. With it, the first detection opens an alert and the following fifty-nine attach to it, leaving one thing to acknowledge.

This article covers how deduplication actually works, using PagerDuty's documented implementation as the concrete example, why the key you choose is the entire design decision, what deduplication cannot do, and how it differs from the grouping and suppression features it is often confused with.

How deduplication works

The mechanism is a key, and it is simpler than the terminology suggests.

Every event sent into an alerting system can carry a deduplication key, which is a string identifying the problem rather than the observation. In PagerDuty's Events API v2 the field is literally named dedup_key. When a trigger event arrives, the system looks for an open alert with a matching key. If it finds one, the new event is appended to that alert's log rather than creating anything new. If it does not, a new alert opens.

Three details from PagerDuty's own documentation determine how this behaves in practice. First, if you omit the key, PagerDuty generates a unique identifier automatically, which means every event creates a new alert and you have effectively disabled deduplication by default. Second, the key is case-sensitive. Third, once an alert is resolved, a subsequent trigger event carrying the same key opens a new alert rather than reviving the old one, while acknowledge and resolve events for a key with no open alert are simply dropped.

There is a scoping rule that catches people out. Subsequent events only attach to the open alert if they arrive through the same routing key as the original trigger. Two events with identical deduplication keys sent to two different integrations on the same service will not deduplicate, and you get two incidents. Our guide to advanced alert configuration covers the surrounding routing and notification design.

The key is the whole design

Call this the deduplication key lifecycle, and the reason it deserves attention is that the key is the only real decision you make. Everything else is mechanics.

A key that is too broad collapses genuinely distinct problems into one alert. If your key is the service name, then a failing database connection and an expired credential in the same service produce a single alert, and whoever acknowledges it sees only the first. The second problem is now invisible, and it stays invisible until somebody resolves the alert and the next event opens a fresh one. This is a common and expensive failure, and it also produces a confusing secondary symptom: after the alert is resolved, new and unrelated events keep arriving under the same key and appear to be a continuation of something already handled.

A key that is too narrow defeats the purpose. Including a timestamp, a request identifier, or anything else that changes between checks means every event is unique and you are back to sixty pages an hour.

The useful shape is a key that names the specific failing thing, stably: the service plus the component plus the check, with nothing in it that varies between observations of the same problem. A nightly build failure for one application, keyed on that application and that pipeline, is a good example. It is stable across repeated failures and distinct from every other pipeline.

Some systems let you override or construct the key with rules after the event arrives, which is useful when the tool sending the events does not let you control the payload. That capability is worth checking for, because it is often the only way to fix a poorly keyed third-party integration.

What deduplication does not do

Deduplication reduces duplicates of one problem. It does not reduce the number of distinct problems, and conflating the two is why teams sometimes turn it on and find their alert volume unchanged.

If a database goes down and thirty services that depend on it each raise their own alert, deduplication does nothing at all, because those are thirty different keys describing thirty genuine observations. Collapsing them is a different capability, usually called alert grouping or correlation, which works on relationships between distinct alerts rather than on repetition of one. It is worth knowing that some of these correlation and suppression capabilities sit on higher pricing tiers in commercial platforms, so check what your plan actually includes before designing around them.

Deduplication also does not decide whether something deserved to be an alert at all. A check that is flapping between healthy and failed every few minutes will, if it fully resolves between failures, open a new alert each time by design. The fix there is a consecutive-failure threshold or a confirmation from a second location before the alert fires, not a deduplication key.

And it does not survive resolution. This is intentional and correct, because a recurrence after a fix is genuinely new information, but it does mean that a problem that keeps coming back produces a sequence of separate alerts rather than one long one. If that pattern is what you care about, you need reporting on alert frequency by key rather than a change to the key itself.

Where the practical wins are

The largest single improvement for most teams is checking whether deduplication keys are being set at all. Because omitting the key silently produces unique identifiers, a custom integration written without reading the documentation will look like it works and will page relentlessly during the first real outage. That is the failure to look for first.

The second is auditing keys that are too broad, which requires looking at resolved incidents rather than open ones. If your alerts routinely contain many events describing several different underlying failures, the key is doing too much work.

The third is aligning keys with ownership. An alert is a unit of human attention, and the ideal key produces alerts that map to one person's remit. When a key spans two teams' systems, whoever acknowledges it inherits a coordination problem along with the alert.

Common mistakes in alert deduplication

Not setting a key at all. The default behaviour generates a unique key per event, which means no deduplication and full-rate paging. Custom integrations built without checking this are the usual source.

Using a key that identifies the event rather than the problem. Anything varying between observations, such as a timestamp or a request identifier, makes every event unique and defeats the mechanism entirely.

Using a key so general it swallows unrelated failures. Keying on the service name alone hides a second, different problem behind the first, and the responder has no way to see it until the alert is resolved.

Assuming deduplication will reduce alert volume during a large outage. It collapses repeats of one problem, not many distinct problems with a shared cause. That is correlation, and it is a different feature that may sit on a different pricing tier.

Forgetting that keys are scoped to the integration. Identical keys sent through two integrations on the same service do not deduplicate. If the same alert is arriving twice, check whether the source is sending through two routes.

FAQ

What is alert deduplication?

It is the mechanism that collapses repeated notifications about the same ongoing problem into one alert. Events carrying the same deduplication key attach to an existing open alert instead of creating new ones, so a failure detected every minute produces one page rather than sixty.

How does a deduplication key work?

It is a string that identifies the problem rather than the observation. When an event arrives, the system checks for an open alert with a matching key and appends to it if one exists. In PagerDuty's Events API v2 the field is called dedup_key, it is case-sensitive, and omitting it causes a unique key to be generated automatically.

Why am I still getting duplicate alerts?

Common causes are no key being set, a key containing something that varies between checks, or the same events arriving through two different integrations. Deduplication is scoped to the routing key, so identical keys sent through separate integrations on one service will not collapse.

What is the difference between deduplication and alert grouping?

Deduplication collapses repeats of a single problem into one alert. Grouping or correlation collapses several distinct alerts that share an underlying cause, such as thirty services alerting because one database failed. They solve different problems and grouping is often a higher-tier feature.

Closing thought

Deduplication is one of the few alerting features where the default behaviour is the dangerous one. A system that generates a fresh key for every event looks entirely healthy right up until something breaks for an hour, and then it produces the exact pattern that teaches people to silence their phones. The fix is not a setting to enable so much as a decision to make carefully, once, about what your key names.

Fewer alerts to deduplicate is better than better deduplication. Odown checks from seventeen global locations at intervals down to one minute on every plan and routes through Slack, Discord, Telegram, Opsgenie, PagerDuty, email, and webhooks, so a genuine failure can be distinguished from a single-region routing blip before it ever becomes an event with a key attached to it.