Skip to main content
InfromatinTechnologies
Engineering6 min read

Explaining an outage to a non-technical board

The technical cause is the least interesting part of the update. A structure that keeps people informed without alarming them.

Infromatin Technologies

Explaining an outage to a non-technical board

During an incident, the technical cause is the least interesting thing you can tell a stakeholder. It is also the first thing people reach for, because it feels concrete.

What a non-technical audience needs is an assessment: is this affecting customers, how much, and when will it be resolved. The root cause can wait for the post-incident review.

A structure that holds under pressure

Four sentences, in this order, updated on a fixed cadence regardless of whether anything has changed:

What is affected. Which journeys, roughly how many users, and whether any data is at risk. Plain language: "customers cannot complete payments at the moment". Not "the payment service is degraded".

What we are doing. The action currently in progress, in one sentence. If the honest answer is "we are investigating", say that.

What we expect. The next checkpoint, with a time. "We will know more by 14:30." The specific time matters more than the content.

What happens next. The follow-up commitment: when the full update will be issued, and by whom.

An update with no change in status is still worth sending. Silence is read as either incompetence or concealment, and neither is true.

Say what you do not know

"I don't know yet, and here is what we are checking" is a credible, professional answer. Teams avoid it because it feels weak, and they say something speculative instead, which turns an incident into a credibility problem.

Mark uncertainty explicitly. It buys credibility for the things you do assert.

Do not speculate on cause mid-incident

It is tempting to explain early, particularly when the cause seems obvious. Resist it. Early causes are frequently wrong, and a retracted explanation during an incident damages trust in everything else you said that day.

Say what is broken in terms of symptoms. Explain the cause afterwards, in writing.

Separate impact from cause when writing it up

The post-incident review should contain two distinct sections, and mixing them causes most of the argument:

Impact — who was affected, for how long, what it cost in transactions, in revenue where estimable, and in support volume. No hypotheses.

Contributing factors — what allowed this, what made detection slow, what made recovery hard. Multiple factors, no blame, and explicitly including the things your own controls permitted.

The impact section is what the board will read. Write it first, and write it with the same care.

Avoid the language that costs you trust

"Users may experience some issues." Vague and slightly evasive. Name the affected journey. "We have identified the root cause." Only say this when you know it. "This is resolved." If monitoring is clean and the fix is deployed, fine. If it is rolled back, say that. "There is no data loss." A serious claim. Make it only if you have verified it, and say how you verified it.

Specificity is what reads as competence, even when the news is bad.

Prepare before you need it

Write the template while things are working. One page: the four-sentence structure, the cadence, the named owner, and the escalation path to whoever needs to make a commercial decision.

During an incident, the person communicating should not also be the person debugging. Splitting those roles is the single highest-value thing you can prepare, and it costs nothing until you need it.

Decide in advance who is authorised to say customer numbers out loud, and who is authorised to commit to a resolution time. Discovering those limits mid-incident is how well-managed incidents become board-level ones.

In this article

  • incident response
  • communication
  • SRE

Working on something similar?

These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.

Start a conversation

Related reading

Continue from here

Articles connected to the same delivery problems.

Have a related problem in front of you?

Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.

We would like to use Google Analytics to understand how this website is used. No analytics are loaded unless you accept. Your choice is stored for six months.

See our Privacy Policy for details.