On-call roster guide: who belongs on it, how many you need, how to keep it fair
A roster is the people side of on-call. Roster vs schedule vs rotation, why three people is the practical floor, a roster template, and rules for load, onboarding and multiple rosters.
Sandeep Sidhu · Founder, AlertKick
An on-call roster is the ordered list of people who take turns being on call for a service or team, plus the rotation rule that decides whose turn it is. It is the people side of on-call: who is in the pool, in what order, and how load is shared between them. The schedule is what the roster produces once a rotation length and a handover time are applied. Three people is the practical floor for a roster that pages out of hours; two is a coverage arrangement that only works until somebody takes a holiday.
Most on-call problems that get blamed on the tooling are roster problems. The wrong people are in the pool, the pool is too small, one person quietly carries more turns than the others, or a new joiner was added to the rotation before they had access to anything. This guide covers the people side. The time side (rotation length, handover time, timezone) is covered in how to create an on-call schedule, and the operating discipline in on-call rotation best practices.
What is an on-call roster?
An on-call roster is the ordered list of people who take turns being on call for a service or team, plus the rotation rule that decides whose turn it is. Other names for the same thing include on-call rota (UK), call roster, and duty roster. In a call centre “roster” means the staffing plan for shifts; this post is about engineering on-call, where the roster feeds a paging tool. The roster answers “who is in the pool and in what order”. It does not contain dates. Dates come from the schedule, which is generated by applying a rotation rule (for example weekly, handover Monday 09:00) to the roster.
In practice the roster lives in the paging tool (AlertKick, PagerDuty, Opsgenie), not in a spreadsheet, because the tool has to know who to page right now and who to escalate to if they do not answer. The next section explains the difference between roster, rotation and schedule in a table.
What is the difference between a roster, a rotation and a schedule?
The roster is who, the rotation is the rule, and the schedule is the calendar that results. Documentation often uses these terms interchangeably, which causes real confusion when a team tries to change one and accidentally changes the others.
| Term | What it answers | Example |
|---|---|---|
| Roster | Who is in the pool, in what order | Priya, Tom, Ana, Wei |
| Rotation | How the pool advances | Weekly, handover Monday 09:00 Europe/London |
| Schedule | Who is on call on which dates | Priya 24 Aug-31 Aug, Tom 31 Aug-7 Sep, … |
| Override | A dated exception to the schedule | Ana covers Tom 2 Sep-4 Sep |
The order matters because it is the only thing that produces a schedule. Add a person to the roster and the schedule regenerates from the next handover. Change the rotation rule and every future shift moves. Overrides sit on top and leave both untouched, which is why editing the roster to cover a holiday is the wrong tool; see roster management for the difference between a swap and a member change.
Who belongs on an on-call roster?
Anyone who can take the first meaningful action on the alerts that roster receives. That means acknowledge, open the runbook, and either fix the problem or know exactly who to escalate to. It does not mean the most senior person, the manager, or everyone on the team.
Use this test for each candidate:
- Do they have production access sufficient to act on the most common alerts, without waking someone else to get it?
- Have they run the runbooks for those alerts at least once during working hours?
- Can they be reached on the notification channels the escalation policy uses, and have they set their own preferences?
- Are they in a timezone where the shift hours are survivable? A person eight hours offset from the rest of the roster on a weekly rotation is doing a very different job from everyone else.
People who fail the first test should not be on the roster, however willing they are. They belong in the escalation chain as a named person or a specialist roster, reached only when the first responder needs them. Putting them in the rotation guarantees that a proportion of alerts go to someone who can only forward them.
Managers are a common special case. A manager who can act is a fine roster member. A manager who cannot is better placed as the last escalation level, the person who gets paged when nobody else has acknowledged, because that is a signal the process has failed rather than a request to fix a server.
How many people should be on an on-call roster?
Three is the floor. Three to six is comfortable. Beyond eight, the roster is usually two rosters that have not been separated yet.
The arithmetic is simple. On a weekly rotation with three people, each person is on call one week in three. With one person on leave, the other two alternate. This is sustainable for a week or two. With a second person unavailable, one person is on call continuously. That is a coverage emergency, not a rotation.
| Roster size | Weekly turns per year | One person away | Two people away |
|---|---|---|---|
| 2 | 26 | Continuous on-call for the other | No cover |
| 3 | 17 | Two alternate | Continuous on-call |
| 4 | 13 | Three rotate normally | Two alternate |
| 5 | 10 | Four rotate normally | Three rotate |
| 6 | 8-9 | Five rotate normally | Four rotate |
| 8 | 6-7 | Seven rotate normally | Six rotate |
The turns column is not the whole story. A person who is on call one week in eight takes their turn rarely enough to have forgotten what changed since last time. Runbooks drift, dashboards move, and the muscle memory that makes a 3 AM page a five-minute job is gone. Keep rosters at six or fewer. Split a large team by area rather than pooling everyone.
What does a two-person roster actually mean?
One of two people is always the primary and the other is always the secondary. There is no third level. Two people is a legitimate starting point for a service that is not yet critical, but treat it as a known risk with a plan to get to three. While it lasts:
- Keep out-of-hours paging to the alerts that genuinely need a human. Everything else waits for working hours. Cutting noise is worth more with two people than with any other roster size.
- Consider a shorter rotation, such as daily or three-day turns, so neither person carries an entire week when the other is unavailable.
- Write down what happens when both are unreachable. If the answer is “nothing”, make sure the people who depend on the service know that.
How do you keep an on-call roster fair?
Fairness in a roster is mostly about the member order and the calendar, not the rotation length. Two things quietly make one person carry more than the others.
The first is fixed handover days combined with public holidays. If handover is Monday 09:00 and the rotation has four people, each person’s turn lands on the same weeks each cycle, and one of them will keep getting the bank holiday weekend. Rotating the member order once a quarter or picking a rotation length that is not a divisor of common holiday spacing spreads it out. The timeline preview that shows the full schedule months ahead in your own timezone is the way to check who has which holiday before it becomes an argument.
The second is overrides that flow in one direction. If one person is always the one who covers, they are effectively on call more than the rotation says. Swaps, which pair two overrides so that both people give and take a shift, keep the ledger balanced. Track overrides per person and look at the count when reviewing the roster.
Load is a separate axis from turns. Two people can have the same number of shifts and very different numbers of pages if their weeks happen to coincide with releases or with a noisy service. Pages per shift, broken down by person, is the number to watch; see what to measure in rotation best practices.
How do you onboard a new engineer onto the roster?
Three stages. Never drop a new engineer straight into the rotation.
- Shadow. The new engineer receives the same notifications as the person on call (a Notify Person level in the escalation policy, or an override that adds them as a secondary) but is not expected to act. They read the runbooks alongside real alerts.
- Reverse shadow. The new engineer takes the primary role for a shift with an experienced engineer as secondary, reachable on a short escalation wait. The experienced engineer does nothing unless asked.
- Solo turn. Add them to the roster. Insert them at a point in the member order that does not give someone else two consecutive turns, and check the preview before saving.
Before stage one, confirm the practical things: production access, VPN, credentials for the monitoring tools, the mobile app installed with push notifications allowed, and their own notification preferences set. Most failed first shifts are access problems, not knowledge problems.
What does an on-call roster template look like?
A roster is a short document. It should live outside the tool that runs it. The fields below are the ones that matter.
| Field | Example | Notes |
|---|---|---|
| Roster name | prod-platform-primary | One per area of expertise |
| Scope | Kubernetes, ingress, CI runners | Which alerts route here |
| Members, in order | Priya, Tom, Ana, Wei | Order produces the schedule |
| Rotation | Weekly, 1 week | Daily, weekly or custom; rotate every N |
| Handover | Monday 09:00 | Working hours, start of the week |
| Timezone | Europe/London | IANA name; the handover is computed in this zone |
| Secondary | prod-platform-secondary or Notify Person: lead | Where escalation goes if primary does not ack |
| Onboarding status | Wei: reverse shadow until 14 Sep | Who is not yet solo |
| Review date | First Monday of each month | When the roster is looked at |
Teams often skip the scope row. Without a written list of which alerts route to this roster, the roster accumulates alerts it cannot act on. Members stop trusting it.
When should a team have more than one roster?
Split when the runbooks and the people who run them differ. One roster per area of expertise is the rule. One roster per service is usually too many. One roster for everything is too few once the team spans more than one specialism.
A five-person platform team that owns the cluster and the CI system almost certainly wants one roster. A team that includes two database specialists and four application engineers, where a replication lag alert would be forwarded by three of the six, wants two. The application roster receives the general alerts. The database roster is either an escalation level below it or the direct target of database alerts via their own escalation policy.
Signs that a roster should be split:
- A recurring alert is routinely acknowledged and then reassigned to the same one or two people.
- Members regularly say they cannot act on a class of alert during their shift.
- The roster has grown past eight and turns are rare.
Signs that two rosters should be merged:
- The same people appear on both, in the same order.
- One of them has fewer than three members and depends on the other for cover anyway.
Multiple rosters per team also cover the primary and secondary pattern. A secondary roster with the same members in a shifted order means the person who was primary last week is never also secondary this week. It gives the escalation policy a Roster level to fall through to before it reaches a named person.
How AlertKick handles rosters
AlertKick rosters are ordered member lists with a rotation type (daily, weekly or custom), a rotate-every-N setting, an IANA timezone and a handover day and time. The editor renders a timeline preview before you save, viewable in your own timezone, so you can check member order and holiday distribution in advance. Overrides and swaps sit on top of the rotation and leave it untouched. iCal feeds exist per roster and per user. Escalation policies can target a roster directly, round-robin through its members, or notify a named person according to their own preferences. Rosters, schedules and escalation policies are included on every plan, including Free. See on-call features and pricing.
Frequently asked questions
- What is an on-call roster?
- An on-call roster is the ordered list of people who take turns being on call for a service or team, together with the rotation rule that decides whose turn it is. The roster is the people side of on-call. The schedule is the calendar the roster produces once a rotation rule and a handover time are applied to it.
- What is the difference between a roster, a rotation and a schedule?
- The roster is who is in the pool and in what order. The rotation is the rule for advancing through that order, such as weekly on Monday at 09:00. The schedule is the resulting calendar of named shifts with dates. Change the roster or the rotation and the schedule regenerates; overrides then adjust individual dates without touching either.
- How many people should be on an on-call roster?
- Three is the practical minimum for a roster that pages outside working hours. With three people on a weekly rotation everyone gets two weeks off for each week on, and one person can be away without the other two alternating indefinitely. Two people is a coverage arrangement rather than a rotation. Between three and six is where most rosters work well; beyond eight, each turn is rare enough that people lose familiarity with the runbooks.
- Can two people run an on-call roster?
- Two people can share a roster, but it should be treated as temporary. Every holiday, illness or resignation collapses it to one person on call continuously, and there is no secondary to escalate to. If two is all you have, keep the rotation short, restrict out-of-hours paging to the alerts that genuinely need a human, and make growing to three a stated goal.
- Who should be on the on-call roster?
- Anyone who can take the first meaningful action on the alerts the roster receives: acknowledge, read the runbook, and either fix the problem or escalate to someone who can. Access, context and permission matter more than seniority. People who cannot act on the alerts should be reached by escalation, not placed in the rotation.
- Should a team have one on-call roster or several?
- One roster per distinct area of expertise, not one per service. If the same people would answer the alert regardless of which service fired it, use one roster. Split when the runbooks and the people who can execute them are genuinely different, for example a database roster and an application roster, and point each escalation policy at the right one.
- How do you add a new engineer to an on-call roster?
- Shadow first, then reverse shadow, then a solo turn with an experienced secondary. Add them to the rotation only once they have the access, the runbooks and at least one supervised shift behind them. Insert them at a point in the member order that does not give someone else two consecutive turns.