Skip to main content
scaling.cloud
Back to the Guide

Following the sun without burning anyone

Round-the-clock coverage does not have to mean round-the-clock suffering. Learn how follow-the-sun rotations, working-hours rules, and humane compensation let a global team hold the pager all night without anyone losing their nights.

There are only two honest ways to cover the small hours. Either someone is awake during their own daytime to take the page, or someone is asleep during their own night-time and you wake them up. Follow-the-sun is the discipline of preferring the first as far as your geography allows, and of being deliberate and generous about the second wherever it does not. A team scattered across enough longitude can chase daylight around the planet and hand the pager off at each timezone's dusk, so that a 3am incident in one region is a mid-afternoon incident for whoever actually holds it. Where you lack the coverage to do that, you owe the people who absorb the nights honesty about the cost and real compensation for paying it.

Hand the pager to wherever it is daytime

The premise of follow-the-sun is simple: at any given hour, route pages to the part of the team for whom that hour is a working hour. A London team carries the European morning and afternoon, hands off to a team in the Americas as their day ends, who hand off in turn to a team in Asia-Pacific, who hand back to London as the sun comes up again. Each region works its own daylight and sleeps its own night, and the pager is never asleep because somewhere it is always day.

This only works if the handoffs are real handoffs, not gaps. A region that signs off at 6pm local without a confirmed, awake successor has not handed off the pager — it has dropped it on the floor for the hours until the next region wakes. Treat the overlap at each boundary as sacred: the outgoing region stays reachable until the incoming region has explicitly acknowledged it holds the pager, and the same short handoff ritual that protects a single-region rotation applies double here, because the context now has to cross not just a shift change but a culture and a continent.

Encode working hours as a rule, not a habit

The mechanism that makes follow-the-sun automatic is a working-hours rule attached to each escalation step: this layer participates only during its region's working hours, and is passed over outside them. The rule carries its own explicit timezone — "UK business hours" means UK time regardless of who or what the step happens to page — because inferring a timezone from anything else is how you end up routing a London team's pages by Sydney's clock. Define each region's window once, in its own zone, and let the rule decide moment by moment whether that region is on the clock.

A subtle but important property: the rule should be evaluated at the instant the page reaches the step, not snapshotted when the incident was first created. An incident that opens at 5:55pm in London and escalates past the primary at 6:10pm should see the London layer as off — its working window closed while the incident was already in flight. Point-in-time evaluation at the moment of the attempt is what keeps follow-the-sun honest across the boundaries it exists to manage.

Skip a closed layer — never hold a page

Here is the discipline that separates a humane follow-the-sun design from a dangerous one. When an escalation step's working-hours rule says "not now," the correct behaviour is to skip that step and immediately try the next one. The wrong behaviour — the tempting, plausible, catastrophic one — is to hold the page until that region's window opens. Holding turns a working-hours rule into a silent delay on an active emergency: the page sits in a queue while a real incident burns, waiting for morning to break somewhere, and the user-facing outage grows by hours for no reason a user would ever accept.

So make the rule a router, never a brake. A region that is off-shift is simply passed over in favour of a region that is on-shift; the page keeps moving down the path at full speed, it just steers around the people who are asleep. If the genuine goal is to suppress low-urgency noise until business hours — and that is a legitimate goal — solve it earlier, by deciding at ingestion that a low-severity signal does not warrant a page at all right now. That is a routing decision about whether to page; it is not the escalation path's job, and bolting it onto the escalation path by holding pages mid-flight is how you teach a region to wake up to a four-hour-old fire.

Pay for the nights you cannot avoid

No amount of geography eliminates every night. Eventually some region's small hours have no daylit successor, and a human there has to take the page from a dead sleep. Pretending otherwise is how good people quietly leave. Be explicit that night coverage is real labour and compensate it as such: time off in lieu of the sleep that was lost, an on-call stipend that scales with how genuinely disruptive the shift is, and a hard cap on how often any one person draws the worst slots. A rotation that pays for its nights can sustain them; one that treats them as an unpaid expectation is borrowing against its people's goodwill at a punishing interest rate, and the balance always comes due.

Coverage and sustainability are the same problem

It is easy to frame follow-the-sun as a trick for buying coverage on the cheap — free overnight monitoring, courtesy of someone else's daytime. That framing curdles fast. The whole point is that coverage and sustainability are not rivals to be traded off but two readings of the same well-built rotation: a schedule that routes each hour to people for whom it is daytime is, by construction, a schedule that lets almost everyone sleep almost every night. Chase the daylight where you can, encode the windows so the routing is automatic, never let a rule become a brake on a live page, and pay honestly for the nights that remain. Do that and the sun never sets on your pager — without anyone on your team having to watch it rise from the wrong side of a 3am incident.

Do this in scaling.cloudSet this up step by step in the docs.Open the how-to guide

Put this into practice.

scaling.cloud gives your team the on-call schedules, paging, and incident timelines this guide describes.