← All posts

The 5 Cs of Effective Outage Communication

View Markdown

Your engineers are already in the incident channel. Support is already fielding "is it just me?" tickets. The question is whether the next message customers see is a clear official update, or a rumor they assembled from Slack, Twitter, and a silent status page.

Customers can tolerate downtime. They struggle to tolerate silence. Research on major-incident response found that 56% of stakeholders were more frustrated by poor communication than by the incident itself, and 87% expect regular status updates until resolution (Everbridge).

The 5 Cs are a writing standard for those updates: clear, consistent, comprehensive, cohesive, and candid. They're not a slogan. They're the difference between a timeline people trust and a timeline people ignore.

Be clear

Weak: "We are currently experiencing an issue with some backend services that may impact certain user workflows."

Better: "Checkout is failing for some customers. Payments are not completing. We're investigating and will post again within 30 minutes."

The second version answers the only questions that matter in the first five minutes: what's affected, what customers should expect, and when you'll speak again. Skip jargon. Skip "degraded performance across distributed subsystems." If a non-technical customer can't repeat the update to their boss, it isn't clear enough.

Lead with what broke, who it hits, and what happens next. Save the root-cause theory for later.

Every first update should include:

  1. What is affected (specific components, not "systems")
  2. What customers should expect (errors, slowness, full unavailability)
  3. What you are doing now
  4. When you will update again

PagerDuty recommends acknowledging a customer-impacting issue within 10 to 15 minutes of detection. You don't need the root cause to be clear. You need the impact to be.

Practical rule: if a customer would still open a ticket after reading the update, you left out the impact or the next update time.

Stay consistent

The fastest way to lose trust mid-incident is to say two different things in two different places.

If the status page says "payment delays," Slack shouldn't say "checkout API errors," and social shouldn't say "intermittent issues." Pick one description of the impact and reuse it across email, SMS, chat, and social. Change the wording only when the situation actually changed.

Treat the status page as the source of truth. Other channels exist to point people there, not to invent a second narrative. PagerDuty's status page guidance treats a dedicated page the same way: one canonical timeline, amplifiers around it.

Consistency is also a staffing problem. Google's incident management model includes a Communications Lead whose job is stakeholder updates, so engineers can focus on mitigation. Name one person who owns the public wording. That prevents duplicate posts, conflicting messages, and gaps during handoffs.

Practical rule: one person owns the public wording per event. Every other channel quotes that wording or links to the status page.

Be comprehensive

"We are aware of the issue" is an acknowledgment, not an update. After the first post, customers want progress: what you tried, what you ruled out, what's still unknown, and when you'll be back.

"Outbound email to some domains is still failing. We confirmed the app is queuing messages correctly and ruled out a DNS change on our side. We're working with our email provider on their outbound relay. We don't have a confirmed cause yet. Next update within 30 minutes."

That's what you tried, what you ruled out, and what you still don't know. Comprehensive doesn't mean long. It means the reader doesn't have to guess what you skipped.

Post on a cadence even when nothing changed. PagerDuty recommends roughly every 30 minutes during major outages. Silence reads as neglect. "No change since the last update. We're still working the fix. Next update in 30 minutes" is better than a quiet hour.

Practical rule: every public update should say what changed (or that nothing changed) and when you will post again.

Keep a cohesive timeline

Customers don't experience your incident as a pile of Slack messages. They experience it as a story: you noticed, you found it, you fixed it, you watched it, you closed it.

A standard incident workflow makes that story readable:

PhaseWhat customers should hear
InvestigatingWe see the impact. Cause isn't confirmed yet.
IdentifiedWe know the cause. A fix is in progress.
MonitoringThe fix is in. We're watching for stability.
ResolvedImpact has ended.

Name the phase in the update. "The cause has been identified and we're implementing a fix. Next update within 30 minutes" tells people where they are in the process. It also tells your team what "done" looks like.

On StatusDashboard, those phases live on the event timeline and follow your incident workflow. Post a new timeline entry when the situation materially changes, or on cadence even if it hasn't. Resolve explicitly. A slow fade back to green without a closing update leaves people wondering whether it's safe to resume work.

Practical rule: if a new reader landed on the event right now, the latest timeline entry plus the current phase should be enough. They shouldn't have to reconstruct the story from five older posts.

Speak with candor

Vague: "An unexpected issue has been resolved and we apologize for any inconvenience."

Better: "Checkout is working again. The outage was caused by a failed configuration change on our payments service, which blocked card authorization for some customers. We'll publish a post-mortem on this status dashboard with the corrective actions we're taking."

The second update names the cause, the impact, and what you'll change. The first one hides. Honesty is the part teams skip when they're embarrassed.

When you know the cause, say it in plain language. When you don't, say that too. When you made a mistake, own it. When you'll publish a post-mortem, commit to it and then actually publish it.

Candor isn't a dump of internal ticket numbers or vendor blame. It's the cause, the impact, and what you'll change so it doesn't happen the same way again.

After significant incidents, publish a post-mortem when you're ready. Resolve the event first. Customers want confirmation that impact has ended before they read the analysis.

Practical rule: if you'd be uncomfortable reading the update as a customer, it's still too vague or too defensive. Rewrite it.

A checklist you can run during the next incident

Before you publish the first update:

  • What is affected, in customer language
  • What customers should expect right now
  • Next update time (even if that is "in 30 minutes")
  • Same wording queued for every channel you will use

On every later update:

  • Current workflow phase
  • What changed since the last post (or "no change")
  • What you are doing now
  • Next update time

When you close:

  • Impact has ended, said plainly
  • Cause, if you know it
  • Whether a post-mortem is coming, and where

StatusDashboard is built around that workflow: draft-first publishing, a public timeline, named phases, and notifications that stay in sync with the page. The habits still belong to your team.

For day-to-day event types and notification habits, see Best Practices for Managing a Status Page. If you're still deciding whether you need a public page, start with Why You Need a Status Page.

Start your free trial and put the 5 Cs in place before the next incident, not during it.

We use cookies

We use essential cookies to keep the site working, and optional analytics cookies to understand how it's used. Read our Privacy Policy.