Site Reliability Engineer Interview Question

What makes a postmortem blameless, and what belongs in one?

What the interviewer is probing, how to structure your answer, and a spoken example you can adapt.

Quick answer

Blameless means the write up asks how the system allowed a reasonable person to take that action, not who broke it. Include a timeline, user impact quantified against the SLO, contributing factors, what went well, and action items with named owners and due dates. Ban counterfactuals such as should have known. The real test is whether the actions actually get completed, so track them like any other work.

Why interviewers ask this

Interviewers are assessing engineering culture instincts as much as process knowledge. They want to hear that blameless is about honest reporting rather than politeness, that impact is measured rather than described, and that action items are prioritized and tracked. Candidates who say we wrote a doc and moved on reveal that nothing changed after their incidents.

How to structure your answer

  • Define blameless as a mechanism for getting accurate information.
  • List the sections that must exist, including what went well.
  • Quantify impact in SLO or business terms, not adjectives.
  • Require owners and dates on every action item.
  • Say how completion is tracked and reviewed.

Example answer

Spoken example, first person

Blameless is not about being nice, it is about getting the truth. If people expect to be punished, you get sanitized timelines and you never learn what actually happened, so the next incident is identical. Concretely I want a timeline with timestamps, quantified impact (we burned 60% of the monthly error budget and 12,000 checkouts failed), contributing factors rather than a single root cause, and a section on what went well, which people skip and which is how you find out that the runbook or the rollback tooling saved you. Action items need an owner and a date, and they go into the normal backlog with real priority, otherwise you have written a nice document that changes nothing. The language rule I enforce is no counterfactuals: not the engineer should have noticed, but the deploy tool let a change go out with no canary and no confirmation. That reframing almost always produces a better fix than a training reminder.

Walking into this interview soon? GhostPilot listens to your live call, spots the question the moment it is asked, and puts a structured answer on your screen in real time. Try it on your next mock, or grab a $29 Session Pass, no subscription, for the real thing.

See how it works

Follow-up questions to expect

  • How do you keep action items from rotting in a backlog?
  • When would you skip a postmortem entirely?
  • How do you handle an incident that was genuinely caused by human error?

Related site reliability engineer questions

Your interviewer will ask their own version of this. Paste your actual job description into the free Question Predictor and get the 20 questions that role is most likely to ask, with what each one is really probing.

Predict my questions

Rehearse the hard questions before they are asked

Practise with a live copilot, then walk in ready. A $29 Session Pass gets you through the interview with no subscription and no lock-in.

Get GhostPilot