How do you run on call without burning the team?

Design a humane rotation that protects service and people

on call rotation works when managers treat it as a service design and people leadership responsibility, not a test of toughness. A rotation becomes harmful when pages, uncertainty, and recovery costs are treated as personal endurance tests. The practical goal is reliable response with fair load, clear support, fewer pages, and genuine recovery. Start with evidence about interruptions and decisions, then change the system before asking individuals to absorb more strain.

Why does this need active management?

On call connects technical reliability with a person's time, sleep, confidence, and family life. A schedule may look tidy in a calendar while the real experience varies sharply. One week can contain no action. Another can bring several pages, uncertain decisions, and hours spent restoring context. Managers need to understand that lived burden because staffing and priority choices shape it.

A six person platform team rotates weekly. Two experienced engineers receive most escalations because alerts lack context and newer engineers don't feel safe making decisions. This is not merely an individual resilience issue. The team's architecture, alert rules, documentation, access, staffing, and delivery pressure all influence what happens. When leaders call the result part of the job, they hide choices that can and should be improved.

Active management doesn't mean joining every response. It means setting expectations, protecting capacity, making backup dependable, and ensuring repeated pain produces planned work. The manager also makes tradeoffs visible to product partners. If reliability work is always displaced by features, the organization is choosing future interruptions even when nobody says so directly.

What outcome should the team design for?

The desired outcome is not simply a filled schedule. A healthy system gives the person on call enough context, authority, tools, and support to make safe decisions. It limits avoidable interruption and recognizes that difficult shifts reduce later capacity. Customers receive reliable service because the team can respond without consuming its people.

Define success in operational and human terms. Service recovery matters, but so do sleep, concentration, confidence, and the distribution of difficult work. A team that restores every incident while repeatedly exhausting two experts isn't healthy. It has borrowed reliability from a small group and will eventually pay through errors, absence, or departure.

Write a simple purpose for the rotation. It might say that on call provides a supported response to urgent customer impact while generating evidence for service improvement. That sentence gives managers a basis for rejecting pages that aren't urgent, training people before solo duty, and protecting repair work after recurring failures.

Which decisions make the biggest difference?

  1. How do you make on call fair? focuses on on call fairness.
  2. Who should be on call? focuses on who should be on call.
  3. What should you do after a hard on call week? focuses on on call recovery.
  4. How can a manager reduce on call pages? focuses on reducing on call pages.
  5. How do you talk about on call in 1 on 1s? focuses on on call 1 on 1.

Measure pages and sleep disruption, clarify escalation, pair experience deliberately, fund alert repair, and protect recovery after difficult weeks. These decisions reinforce one another. Better readiness reduces avoidable escalation. Fewer pages make recovery more realistic. Fair participation spreads knowledge, and private conversations reveal costs that dashboards miss. Treating only one part usually moves the burden rather than removing it.

Managers should distinguish policy from judgment. A policy can define shift length, response expectations, backup, and compensation. Judgment is still needed when a difficult week, health need, unfamiliar service, or launch changes the risk. Consistency means using the same principles, not forcing every person through an identical experience.

What evidence should a manager review?

Begin with page volume, interrupted sleep, escalation concentration, recovery time, and whether engineers can explain the service. Review trends by service, shift, and type of interruption. Totals alone can conceal concentration. Ten pages spread across ten people aren't the same as ten pages waking one person. Add context about duration and difficulty because a short acknowledgement differs from a long uncertain recovery.

Use conversation alongside operational data. Paging systems show timestamps, not fear before a shift, lost sleep after resolution, or the effort of asking a reluctant expert for help. Ask engineers what felt unclear, which pages required action, where they hesitated, and what would have made the week safer. Their account is evidence, not a complaint to discount.

Keep measurement focused on system improvement. Don't rank engineers by response speed or number of pages closed. Hard incidents and careful decisions take longer. Individual scoring encourages people to act alone, delay escalation, or hide uncertainty. Look for patterns that guide staffing, training, alert repair, and ownership.

How should a manager respond this week?

Choose one recent rotation and reconstruct the experience with the engineer. Review each meaningful interruption, the decision required, available context, backup response, and effect on later work. Ask what should be removed, clarified, practiced, or repaired. Then select a small action with an owner and a date rather than producing a broad promise.

Adjust current commitments when the evidence shows real cost. Reliability work and recovery need calendar space. Saying they are important while preserving every feature date leaves the tradeoff with the engineer. A manager should tell partners which work moves and why. This makes operational health a leadership choice instead of invisible personal labor.

Share relevant learning without exposing private details. The team should know which alert will change, which runbook needs work, or which backup rule is now clearer. Personal constraints and emotional disclosures stay private unless the employee wants otherwise. Trust grows when managers convert information into action with appropriate boundaries.

How do you keep the approach fair?

Fairness begins with impact, not identical treatment. People have different service knowledge, health needs, caring duties, locations, and access to quiet recovery time. Managers can preserve a common responsibility while making reasonable adjustments. The key is to discuss constraints privately, explain principles clearly, and prevent accommodations from quietly reducing opportunity.

Recognition must include prevention and support. The person who resolves a dramatic page is visible, while the engineer who removes a noisy alert or patiently backs up a new colleague may be overlooked. Performance conversations should value service repair, teaching, documentation, and sound escalation. Otherwise the reward system favors heroics over a stable team.

Review who receives the hardest services, overnight interruptions, and repeated backup requests. Also review who gets protected project work after a difficult shift. If burden and opportunity follow the same people repeatedly, change assignment and invest in broader capability. Transparency about the pattern is more useful than insisting the schedule is neutral.

What should happen when someone needs help?

Backup must be a real operating role. The primary engineer should know whom to contact, how quickly help should arrive, and which decisions require escalation. Managers need to reinforce that early help is sound judgment, not weakness. A person who waits because they fear looking unprepared can turn a manageable event into customer harm.

Practice common decisions before a live shift. Walk through a failing dependency, unclear alert, rollback choice, and customer communication. Let the learner use actual tools while the experienced engineer observes. Confidence built through rehearsal is more dependable than confidence requested through encouragement alone.

After support is used, ask whether the gap was expected. Some escalation is appropriate for rare complexity. Repeated escalation for the same service points to missing knowledge, access, documentation, or design. Assign a correction rather than continuing to depend on informal favors from the same expert.

Which management mistake causes the most harm?

The central mistake is celebrating endurance while leaving noisy systems and uneven expertise unchanged. It sends the message that the individual must adapt to a system leadership chose not to improve. People may comply for a while, especially when reliability work is associated with seniority or commitment, but silent compliance isn't proof that the load is safe.

Another mistake is reacting only after a crisis. Small repeated pages, mild dread, and frequent backup requests are leading signals. Review them before they become absence, conflict, a serious error, or resignation. Regular attention lets managers use smaller changes and gives the team evidence that speaking early is worthwhile.

Avoid promises without ownership. Saying that alerts should improve or documentation should be better doesn't alter anyone's next shift. Name the person who will act, the time available, and the review date. If roadmap pressure blocks the work, record that decision and revisit the risk with the leader who owns the priority.

How can leaders protect long term capability?

Build service knowledge during normal hours. Pair engineers on changes, rotate operational reviews, improve runbooks through use, and let people observe experienced decisions before carrying responsibility alone. A rotation shouldn't be the first place someone discovers how a critical service behaves. Prepared learning lowers fear and reduces dependence on memory.

Reserve capacity for prevention every planning cycle. Use page evidence to select work with a clear expected effect, such as removing a false alert, simplifying diagnosis, or fixing a repeated cause. Review whether the change reduced burden. Reliability investment becomes credible when the team can connect planned effort with quieter shifts.

Protect the feedback loop in 1 on 1s and team reviews. Ask what changed, what remains hard, and whether prior commitments happened. Engineers notice quickly when leaders collect concerns but don't return. Closing the loop is part of the intervention because it determines whether people will share the next weak signal.

What should the manager do next?

Start with the next scheduled shift, not a complete program. Confirm readiness, backup, service context, and current alert risks. Check recent burden and any personal constraint that affects the plan. Make one concrete improvement before the shift and one commitment to review the experience afterward.

Then place a recurring review on the team's operating calendar. Examine the human and technical evidence together, choose prevention work, and verify completion. Keep the discussion blameless but specific. The aim is not to prove that on call is easy. It is to make difficult responsibility supported, bounded, and capable of improving.

on call rotation becomes sustainable when managers treat every rotation as information about the system. Listen to the person, inspect the service, fund the repair, and adjust the work. That rhythm protects customer reliability while showing engineers that their health and judgment matter as much as the appearance of a complete schedule.

What do managers often ask?

How do you run on call without burning the team?

Measure pages and sleep disruption, clarify escalation, pair experience deliberately, fund alert repair, and protect recovery after difficult weeks. Review both service outcomes and the human cost, then give every improvement a clear owner and date.

Which evidence matters most?

Review page volume, interrupted sleep, escalation concentration, recovery time, and whether engineers can explain the service. Combine operational records with private conversation because dashboards cannot show the full personal cost.

How often should managers review the rotation?

Review it after any difficult shift and on a regular monthly rhythm. Check whether promised alert, training, staffing, and recovery actions actually happened.

Should every engineer have the same on call load?

Not necessarily. Fairness considers actual burden, readiness, service difficulty, personal constraints, support, and opportunity rather than calendar turns alone.

What should happen when on call harms health?

Reduce exposure promptly, support recovery, discuss appropriate workplace support, and change the system that created the harm. Don't treat health impact as a toughness problem.

Related: on call fairness, who should be on call, on call recovery.

Prefer prepared conversations over memory alone? Explore iSilta features or try the product demo.