Restore service first, investigate second. Get the fastest safe mitigation in place, usually a rollback unless rolling back is riskier, keep a clear incident commander and a separate communications owner, and protect the engineer involved from having to fix and explain at the same time. Afterward, run a blameless review that finds the system weaknesses (missing guardrails, unsafe deploy path, thin tests) rather than the person.
Why interviewers ask this
Interviewers are checking composure under pressure and your postmortem culture in one question. They want restore before diagnose, defined roles during the incident, honest customer communication, and a genuinely blameless review with tracked actions. The real tell is how you talk about the engineer who shipped the change, because a manager who lets blame land on an individual ends up with a team that hides problems.
How to structure your answer
- Restore service before you diagnose the cause.
- Name the roles: commander, communicator, hands on keyboard.
- Shield the engineer involved from the blast radius.
- Run a blameless review with owned, tracked actions.
Example answer
Mitigation first. The only question in the first five minutes is what is the fastest safe way back to a working state, and usually that is a rollback rather than an inspired fix under pressure. I make sure somebody is incident commander and somebody else owns communication, because the worst incidents I have seen were ones where the person typing was also answering the executive who kept joining the call. If the engineer who shipped the change is the best person to fix it, they fix it, and I take every other conversation off their plate. Once it is stable, the review is about the system. When this happened on my team, a config change bypassed staging because the deploy path allowed it. The engineer was the last domino, not the cause. The actions were a guardrail on that path, a test for that specific failure, and better alerting, each with a name and a date. And I say out loud in the review that the change was reasonable given the information they had, because the whole team is watching how that person gets treated.
Walking into this interview soon? GhostPilot listens to your live call, spots the question the moment it is asked, and puts a structured answer on your screen in real time. Try it on your next mock, or grab a $29 Session Pass, no subscription, for the real thing.
See how it worksFollow-up questions to expect
- How do you decide when to tell customers?
- What if the same engineer causes a second incident a month later?
- How do you make sure postmortem actions actually get done?
Related engineering manager questions
Your interviewer will ask their own version of this. Paste your actual job description into the free Question Predictor and get the 20 questions that role is most likely to ask, with what each one is really probing.
Predict my questions