Short answer
Incident communication follows four standard stages: Investigating (you know something is wrong), Identified (you know the cause), Monitoring (a fix is in place) and Resolved (service is back to normal). Each update should state the affected component, the user impact, what you are doing and when the next update will come. After a significant incident, publish a blameless postmortem covering timeline, root cause, impact and follow-up actions.
Incident communication templates are pre-written messages for each stage of an outage, so the team can post accurate updates in seconds instead of drafting under pressure. The templates below are designed for status pages, but work equally well in email, in-app banners and social posts.
The four incident stages
| Stage | Meaning | When to post |
|---|---|---|
| Investigating | A problem is confirmed; cause unknown | Within minutes of confirmation |
| Identified | The cause is known; a fix is in progress | As soon as the cause is clear |
| Monitoring | A fix is deployed; watching for recovery | After the fix is applied |
| Resolved | Service is fully restored | After a stable period with no errors |
Not every incident passes through every stage. A short blip may go straight from Investigating to Resolved; that is fine.
Principles for every update
- Lead with impact. Say what users experience before anything else.
- Name the component and scope. "API requests in the EU region" is more useful than "some issues".
- Use plain language. No internal hostnames, ticket numbers or jargon.
- Timestamp with a time zone, or use UTC consistently.
- Promise the next update and keep the promise, even if the update is "still investigating".
- Do not speculate or assign blame, including to vendors, until you are certain.
Investigating templates
We are investigating an issue causing [errors / slow responses / failed logins] on [component] for [all / some] users. Our team is working on it, and we will share an update within [30] minutes.
We are aware that some customers are unable to [action, e.g. complete checkout]. We are investigating and will update this page by [time, time zone].
Identified templates
We have identified the cause of [symptom] affecting [component]. A [configuration change / database issue / problem with a third-party provider] is preventing [action]. We are [rolling back / applying a fix] now. Next update within [30] minutes.
The issue has been traced to [plain-language cause]. Data is not affected. We expect to restore service within approximately [estimate] and will confirm when complete.
Only include an estimate if you are reasonably confident. A missed estimate damages trust more than none.
Monitoring templates
A fix has been deployed and [component] is recovering. Error rates are returning to normal. We are monitoring closely and will mark this incident resolved once we confirm stability.
Service has been restored. Some [requests / emails / jobs] queued during the incident may take up to [duration] to process. We are monitoring the backlog.
Resolved templates
This incident has been resolved. From [start time] to [end time, time zone], [component] was [unavailable / degraded] for [scope]. The cause was [one-sentence cause]. We have [immediate fix] and will [preventive step]. We apologize for the disruption.
Resolved. All systems are operating normally. A full postmortem will be published within [5 business days].
Degraded performance template
[Component] is responding more slowly than usual. Most requests succeed, but some may take up to [duration] or time out. We are investigating and will update within [30] minutes.
Scheduled maintenance templates
Scheduled: We will perform scheduled maintenance on [component] on [date] from [start] to [end, time zone]. During this window, [expected impact, e.g. the dashboard will be read-only]. No action is required.
In progress: Scheduled maintenance on [component] is under way. We expect to finish by [time, time zone].
Completed: Scheduled maintenance is complete and all systems are operating normally. Thank you for your patience.
Third-party outage template
An outage at one of our infrastructure providers is affecting [component]. We are in contact with the provider and working on mitigation. We will update within [30] minutes.
Update cadence
| Severity | Example | Update every |
|---|---|---|
| Major outage | Product unavailable for all users | 15 to 30 minutes |
| Partial outage | One feature or region down | 30 to 60 minutes |
| Degraded performance | Slow but working | 60 minutes |
| Maintenance | Planned work | Start, any change, completion |
These are common starting points rather than a standard; adjust to your users' expectations.
Who writes the updates
Assign one person to communications so the engineers fixing the problem can focus. In small teams, the person on call often does both, which is why templates matter: they reduce writing to filling in blanks. In larger teams, a common split is:
- Incident lead: coordinates the response and decides when stages change.
- Communications lead: posts status page updates, keeps support and customer-facing teams informed, and tracks when the next update is due.
- Responders: investigate and fix, and report findings to the incident lead.
Agree in advance who may post publicly and whether updates need a second pair of eyes. A quick review catches wrong times, internal jargon and speculation before they reach customers.
Postmortem outline
A postmortem is a written review of a significant incident that explains what happened and what will change. Keep it blameless: focus on systems and processes, not individuals. A practical outline:
- Summary. Two or three sentences: what broke, for how long, who was affected.
- Impact. Duration, affected components, percentage of users or requests, any data impact.
- Timeline. Timestamped events, from the triggering change through detection, response and resolution.
- Detection. How the problem was noticed, by an alert or by a customer, and how long it took.
- Root cause and contributing factors. The technical cause and the conditions that allowed it.
- Resolution. What fixed it.
- What went well and what did not.
- Action items. Specific, owned, dated follow-ups that reduce the chance or impact of recurrence.
A shortened public version of the postmortem, without internal details, can be linked from the resolved status page update.
Using templates in Uptime Tracker
In Uptime Tracker, incidents open automatically when downtime is confirmed. From the incident you can acknowledge, assign an owner, attach a runbook link, keep internal notes and post public updates that appear on your status page and go to email subscribers. Keep these templates in your runbook so whoever is on call can paste and adapt them. See incident management and what is a status page.
Frequently asked questions
What are the stages of incident communication?
The standard stages are Investigating (problem confirmed, cause unknown), Identified (cause found, fix in progress), Monitoring (fix deployed, watching recovery) and Resolved (service fully restored). Short incidents may skip directly from Investigating to Resolved.
What should the first incident update say?
The first update should confirm that you are aware of the problem, name the affected component and the user impact, and say when the next update will come. It should be posted within minutes, even before the cause is known.
How do you write an incident resolved message?
State that the incident is resolved, give the start and end times with a time zone, describe the impact and scope, summarize the cause in one sentence, mention what you are doing to prevent recurrence, and apologize briefly. Link to a postmortem if one will follow.
What is a blameless postmortem?
A blameless postmortem is an incident review that focuses on how systems and processes allowed a failure, rather than on who made a mistake. It includes a summary, impact, timeline, root cause, resolution and owned action items.
Should status page updates mention third-party providers by name?
Usually not during the incident. Say an infrastructure provider is affected and focus on your users' impact and your mitigation. Naming a vendor before the cause is confirmed risks being wrong and shifts focus away from what customers need to know.