Coach Through Failure — Mentoring Engineers, Part 6
Mentoring Engineers · Part 6
A second-year engineer merges a migration on a Thursday afternoon. It adds a non-null column with a default to a table holding about forty million rows. On the staging database, which has a few thousand rows, it finished in under a second. In production it takes a lock that holds writes on the orders table for eleven minutes. Checkout fails. The incident channel fills up. By the time it's over, the engineer has posted "I'm so sorry, this was me" three times and gone quiet.
What happens to that engineer over the next 48 hours decides two things: whether they learn what actually went wrong, and whether the rest of the team learns that it's safe to say "this was me" out loud. Most mentors handle the first part reasonably well and the second part by accident.
This part of the path is about doing both on purpose. The working claim: a failure is only a lesson if the investigation reaches the decision the engineer made and the information they had when they made it. Stop at "be more careful" and you've taught nothing except fear.
Before you start
Find your team's last three incidents or significant bugs, and the write-ups if they exist. You'll audit them in Step 3.
Know who on your team has caused a production issue in the last six months. If the honest answer is "nobody", that's information too (Step 6).
Read your current postmortem template, if you have one. The structure of the template quietly decides what questions get asked.
Step 1: Separate the incident from the coaching (10 min)
During the incident, the engineer is not your mentee. They're a responder, and probably the one with the most context. Your job is to keep them useful, not to process how they feel. A single private message does more than any speech: "You're not in trouble. Stay on the channel, tell us what the migration does, and we'll talk properly tomorrow."
The fork here is timing. The instinct is to reassure at length in the moment, or to ask "what happened?" while the system is still down. Both pull the engineer into their own head at the point where the team needs their knowledge of the change. Hold the coaching until the system is stable and they've slept.
Test you can run: in your last incident, how long after resolution did someone talk one-to-one with the engineer closest to the cause? Under 24 hours with a calm, scheduled conversation is the target. "Never, it was obviously fine" usually means someone spent a weekend replaying it.
Step 2: Walk the decision, not the outcome (20 min)
Here is the difference in one exchange. The outcome conversation goes: "You should have tested the migration against production-sized data." True, useless, and it lands as a verdict.
The decision conversation reconstructs what the engineer saw at the moment they clicked merge:
Question | What it surfaces |
What did you check before merging? | The migration ran on staging, and CI passed. |
What would have told you the table was large? | Nothing in the PR. Row counts live in a dashboard they'd never been shown. |
Did anyone review the migration? | Yes, a senior approved it. They also didn't flag the lock. |
Was there a rule about migrations on large tables? | Informally, "talk to Priya first". It isn't written anywhere. |
That table is the coaching. The engineer learns the mechanism: adding a column with a default rewrites or locks the table in some database engines and versions. They also learn the discovery move: before changing a table, ask how big it is and how hot it is. The team learns that the safeguard it believed it had, a reviewer plus an unwritten rule, was never a safeguard.
This is the stance John Allspaw described at Etsy in 2012 as a just culture: you investigate the situation and the decisions of the people close to the failure, so that they can give a full account without fearing punishment. The point isn't to be kind. It's that a punished engineer gives you a short account, and a short account hides the system's gaps.
Test you can run: reread your last postmortem. Does it name what information the person had at the moment of the decision? If the "root cause" line names a person, or a quality like "lack of attention", the investigation stopped too early.
Step 3: Let them own the write-up, with a structure that forces depth (15 min)
Have the engineer draft the postmortem, and sit with them for the first pass. Ownership matters: writing it is how they process it, and it shows the team they aren't being hidden. But don't hand over a blank page. Give them three questions that can't be answered with "I'll be more careful":
What signal existed that could have warned us, and where did it live?
What made the risky action look normal at the time?
What change would make this mistake hard to repeat, for anyone?
The third question is the one that matters. In the example, the answers might be a CI check that flags migrations touching tables over a size threshold, a row-count column added to the schema docs, and "talk to Priya" written into the migration checklist. Those are controls. "Be more careful" is a hope.
Google's SRE book (2016) makes the same point about postmortem culture: the write-up exists to find and fix contributing causes, and it should be reviewed and shared, not filed away. Our postmortem template that encourages honesty is built around these questions.
Test you can run: count the action items from your last five postmortems that changed a system, a check or a document, versus the ones that asked a person to try harder. If most are "try harder", your process is producing guilt, not learning.
Step 4: Keep the bar and add support, for a fixed time (10 min)
The quietest damage after a failure is the reassignment. The engineer who broke checkout gets nudged toward documentation tickets and front-end copy changes for a month. Nobody says why. They understand exactly why.
The alternative is to keep the work at the same level and make the extra support explicit and time-boxed: "Next migration, we pair on it. The one after that, you lead and I review. After that, normal process." Writing down the end date is what stops temporary extra scrutiny from turning into a permanent label.
There's a real exception. If the failure exposed a genuine skill gap, something they were never taught, then a short, explicit learning detour is fine. Name it as that: "Let's spend a week on how our database locks work, then you're back on the schema work." The test is whether you can say the reason out loud.
Test you can run: look at what the engineer was assigned in the four weeks after their last significant mistake. Compare it with the four weeks before. If the complexity dropped and nobody can say why, you've lowered the bar without meaning to.
Step 5: Treat a repeat as a question about the system first (10 min)
A second similar failure feels different, and it should prompt a more direct conversation. But walk the same decision questions before deciding it's a person problem. Repeats come from three places, and they need different responses:
What a repeat usually means | What it looks like | The response |
The control from last time was never built | The action item is still open | Fix the system, and own that the delay was the team's |
A skill gap that wasn't addressed | They can't explain the mechanism when asked | A focused learning plan with a check-in date |
Judgment or attention under pressure | They knew, but skipped it to hit a deadline | A direct conversation about trade-offs, and about who set that deadline |
Only the third row is mostly about the individual, and even then the pressure usually came from somewhere.
Test you can run: for any repeat failure on your team, check whether the action item from the first one was completed. In most teams, a good share of "repeat mistakes" are really "repeat open action items".
Step 6: Make failure visible from the top (10 min)
Engineers calibrate how dangerous a mistake is by watching what happens to other people's mistakes, especially senior ones. If the lead's own failures are never mentioned, a junior's first outage feels like a career event.
Talk about your own failures specifically, not as war stories. A useful pattern: "When I dropped the wrong index in 2023, the signal I missed was the query plan, and what we changed afterwards was the review checklist." It has the same shape as a postmortem, so it models the habit as well as the attitude.
Amy Edmondson's 1999 study of 51 work teams found that team psychological safety, a shared belief that the team is safe for interpersonal risk-taking, was associated with learning behaviour such as asking for help, seeking feedback and discussing errors. The mentoring version is simple: people report what they believe is safe to report.
Test you can run: in the last quarter, how many incidents or near-misses were raised by the person who caused them, rather than discovered by someone else? If most are discovered, people are hiding mistakes until they're found.
Worked example: the eleven-minute lock (composite scenario)
Context. A payments-adjacent team of seven, including two engineers in their first two years. Schema changes go through PR review and CI. No size-aware checks exist.
What happened. The migration from the opening locked the orders table for eleven minutes. The engineer apologised repeatedly in the incident channel and then stopped talking.
The mentor's moves. During the incident: one private message, "you're not in trouble, stay on and explain the migration". The next morning: a 30-minute conversation walking the decision table from Step 2. The engineer drafted the postmortem with the three questions from Step 3. The senior reviewer co-signed it and added their own miss: they approved without checking the table size either.
What changed. Three action items, all controls: a CI check that warns on migrations against tables above a set row count, row counts added to the schema docs, and a written migration checklist. The engineer led the next large-table migration three weeks later, paired with the reviewer, using an online schema-change approach.
Evidence it worked. Two months later, a different junior posted in the team channel before merging: "This touches the events table, which the docs say is large. Can someone look at the lock behaviour?" That message is the outcome you're coaching for. It isn't the absence of failure. It's the presence of a question asked early.
Lesson. The engineer learned how database locks work. The team learned that its real safeguard had been one person's memory. Both lessons came from walking the decision, not the outcome.
Trade-offs to make deliberately
Choice | You gain | You risk |
Engineer writes the postmortem | Ownership, deeper processing | It turns into a confession if you don't structure it |
Mentor writes it | Speed, an even tone | The engineer never works through the mechanism |
Public, named postmortem | Team-wide learning, visible safety | Exposure if your culture isn't ready yet |
Anonymised postmortem | Protects the individual | Signals that naming is dangerous, which is the opposite of the goal |
There's no universal right answer on naming. Where trust is high, named and blameless works well. Where it's low, start anonymised and move toward named as the team watches nothing bad happen.
Common failure modes
Reassuring at length during the incident. It takes the engineer's attention off the system when the team needs it most.
Stopping at "be more careful". It produces guilt and no controls, so the same failure comes back with a different name.
The silent reassignment. Easier work with no explanation reads as lost trust, and the engineer will be right to read it that way.
Treating every repeat as a character flaw. Check the action items first.
A lead whose mistakes are invisible. It sets an impossible bar and teaches people to hide.
Final takeaway
Failure becomes a lesson when the investigation reaches the moment of decision and asks what the engineer could see from there. Do that, and the engineer learns the mechanism, the team finds its missing controls, and everyone watching learns that saying "this was me" early is safe. Your next action: take the last postmortem your team wrote and add one row that names the information the person had when they made the call. If you can't fill it in, schedule the 30-minute conversation that would.
Sources
Allspaw, J. (2012). Blameless PostMortems and a Just Culture. Code as Craft, Etsy Engineering.
Edmondson, A. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly, 44(2), 350–383.
Beyer, B., Jones, C., Petoff, J. and Murphy, N. R. (eds.) (2016). Site Reliability Engineering, Chapter 15: Postmortem Culture: Learning from Failure. O'Reilly / Google.
Continue the Mentoring Engineers path
Previous — Part 5: Sponsor, Not Just Mentor
Part of the Mentoring Engineers learning path. Related: From Blame to Learning: Incident Analysis That Drives Change.


