How Big Should an On-Call Rotation Be?

Farouk Ben. - Founder at OdownFarouk Ben.()
How Big Should an On-Call Rotation Be? - Odown - uptime monitoring and status page

How Big Should an On-Call Rotation Be? The Page Budget

Six to eight people is the figure most teams converge on for a weekly rotation, which puts each engineer on call roughly one week in six to one week in eight. That number is a convention rather than a standard, and it is a proxy for the thing that actually determines sustainability: how many pages a person absorbs per shift, and how many of those arrive at night.

This article covers why the six-to-eight range exists, the arithmetic that should replace it, what to do when your team is genuinely too small, and how rotation length interacts with size. The important reframing is that rotation size is a symptom. A rotation that hurts at six people usually has an alerting problem rather than a headcount problem.

Where six to eight comes from

The reasoning is recovery time rather than fairness. With a weekly handoff, a six-person rotation means five weeks between shifts. That is long enough that a bad week has faded before the next one starts, and short enough that people stay familiar with the systems and the tooling.

Below about five, two things break at once. The gap between shifts shrinks to a month or less, so a run of bad weeks compounds instead of resolving, and any absence has an outsized effect. In a four-person rotation, one person on holiday and one on parental leave leaves two people alternating weeks indefinitely, which is not a rotation. Above about ten, the opposite problem appears: shifts become rare enough that people arrive rusty, having not seen the runbooks or the escalation tooling since the previous quarter.

None of this is a rule, and it is worth being precise about that. It is a range that a lot of teams arrived at independently, which is meaningful evidence, but it is not published by a standards body and it does not account for how noisy your alerts are. Our guide to on-call rotations covers the broader scheduling and handoff practices that sit around the number.

The page budget, which is the number that matters

Call this the page budget: the count of pages a person receives per shift, split by whether they arrived during working hours or outside them. It is the metric to manage, and rotation size is one of several levers that move it.

The generally healthy target most teams aim for is no more than a couple of pages per shift outside working hours, and ideally zero on a normal week. Once out-of-hours pages become routine, you are not running an on-call rotation, you are running a night shift without acknowledging it, and adding people to the rotation spreads that load without reducing it.

Work the arithmetic in the direction that exposes the real problem, because it is not the direction most people run it. If your service generates forty out-of-hours pages a month, a weekly rotation means whoever is on call absorbs about nine of them per shift, and that number does not change no matter how many people are in the rotation. Adding people changes how often each person takes the shift, not how bad the shift is. Going from six to twelve halves the shifts each person works in a year, from roughly nine to roughly four, and leaves every individual week exactly as punishing as it was. Halving the alert volume achieves more than doubling the rotation, and it is usually cheaper. Alert deduplication is often the fastest single reduction available. That is the analysis to do before hiring or merging teams for on-call reasons.

The corollary is a useful diagnostic. If the same person is consistently getting more pages than the rest, the problem is routing or an unbalanced service ownership map, not the schedule.

When the team is genuinely too small

Plenty of teams cannot reach six people, and pretending otherwise is not useful. There are four real options and one that only looks like one.

Combine teams for the rotation. Two three-person teams that own related services can run a single six-person rotation, provided both sets of runbooks are good enough that a responder can act on a service they do not own. This is the most common answer and its viability is entirely determined by documentation quality.

Reduce what pages. Narrow the set of alerts that wake someone to genuine customer-facing failures and route everything else to a queue reviewed in the morning. Most small teams discover that a large share of their night pages did not require a night response.

Follow the sun. If you have people in more than one region, aligning rotations to daylight hours removes night pages entirely for most of the week. This requires real coverage in each region rather than one person in a distant timezone.

Accept a smaller rotation and compensate explicitly. A four-person rotation can work if the page volume is genuinely low. Make the trade visible, pay for it, and give time back after bad weeks.

The option that is not one is a rotation of two with an informal expectation that the other person helps out. That is a rotation of one with extra ambiguity, and it burns people quickly.

How rotation length interacts with size

Weekly is the default and it usually deserves to be, because it minimises handoffs, and every handoff is an opportunity to lose context. But length and size trade against each other, and the trade is worth making deliberately.

Shorter shifts of two to three days reduce the damage a single bad shift does and suit small teams with noisy services, at the cost of more frequent handoffs and more people needing current context. Split shifts, where days and nights are covered by different people, are worth considering once night volume is high enough to disrupt sleep regularly, though they need roughly twice the headcount to sustain.

Whatever the shape, the handoff carries the load. A shift report at the end of every rotation, including quiet ones, covering active incidents and their state, anything currently silenced and when the silence expires, and any risky changes landing during the incoming shift, is what keeps a rotation from resetting to zero context every week.

Common mistakes in sizing an on-call rotation

Treating headcount as the fix for a noisy rotation. Adding people spreads the shifts around without making any single shift quieter. Cut the alerts first and size the rotation afterwards.

Sizing for the average week. Rotations are judged on their worst weeks. A rotation that is comfortable most of the time and brutal during a launch will lose people based on the launch.

Counting people who are not really available. A rotation of eight where two are usually travelling and one is new is a rotation of five. Count only people who can genuinely respond.

Ignoring the distribution of pages across people. If one person consistently gets paged more than the others, the schedule is not the problem. Look at alert routing and at which services are attached to which rotation.

Skipping the handoff on quiet weeks. The quiet weeks are when the habit is built. A team that only writes a shift report after a bad week will not have one when it matters most.

FAQ

How big should an on-call rotation be?

Six to eight people is the common convention for a weekly rotation, giving each person a shift roughly every six to eight weeks. Below five, absences make the schedule unworkable; above ten, people lose familiarity with the tooling between shifts.

What is the minimum viable on-call rotation size?

Four is workable only if page volume is genuinely low and the trade is acknowledged and compensated. Below that you are relying on individuals rather than a rotation, and a single holiday breaks the schedule.

Is it better to add people to the rotation or reduce alerts?

Reduce alerts, almost always. Adding people halves how often each person is on call but leaves every individual shift exactly as noisy, while doubling the number of people who must stay current on the systems. Halving alert volume improves the shift itself.

How long should each on-call shift be?

Weekly is the default because it minimises handoffs, and every handoff risks losing context. Shifts of two to three days reduce the damage from a single bad week and suit small teams with noisy services, at the cost of more frequent transitions.

Closing thought

The question is usually asked as a staffing question and is almost always a signal question. Teams that ask how many people they need for on-call are generally asking because the current rotation hurts, and rotation size is the most visible variable. It is rarely the one doing the damage. Count the pages first, split them by hour, and the answer usually turns out to be about which alerts deserve to wake somebody.

Getting that count down starts with alerts that fire on something real. Odown checks from seventeen global locations at intervals down to one minute on every plan, and confirming a failure from multiple regions before it becomes a page is what keeps a routing blip from consuming one of your rotation's night budget.