Skip to content
QSWEQB

Engineering management fundamentals

The answers an EM loop is built from: what changes when you stop being the best engineer in the room, feedback that lands, what a level actually means, how a forecast is stated honestly, and where metrics get gamed. Fifty-eight items, fifteen worked through with an agenda, a ladder, a script or a diagram.

58 questions

Go deeper on Engineering Management

The transition to managing

What actually changes when you stop being the best engineer in the room?

The unit of work changes from a change you made to a change the team made, and with it the feedback loop. As an engineer you knew by Thursday whether the week had gone well because something either worked or did not; as a manager the signal arrives in weeks and is second-hand, which is why new managers feel unproductive while being busier than ever. The second change is that your technical opinion now carries positional weight, so an idea you offered as a suggestion is heard as an instruction and the room stops arguing with you. That is a real loss of information, and the practical response is to speak last, ask what the objections are, and make it explicit when you are thinking aloud rather than deciding.

Show me a one-to-one agenda with what belongs in each part.

Candidates say they run weekly one-to-ones and are rarely able to say what is in one, so the shape is worth having ready.

Weekly, 30 minutes, same slot, recurring. Cancelled by them, never by you.

part                 minutes  owner   what belongs here
-------------------  -------  ------  --------------------------------------
Their agenda first      12    them    whatever they brought. Blockers,
                                      frustrations, a decision they want
                                      cover for, something personal.
                                      If they bring nothing, ask about the
                                      week rather than filling it yourself.

Feedback both ways       6    shared  one specific observation from the
                                      week, given in situation-behaviour-
                                      impact form. Then: what should I do
                                      differently.

Growth and direction     8    you     progress against the thing they said
                                      they wanted. Reference the ladder,
                                      name evidence you have seen, name
                                      the gap. Monthly depth, weekly touch.

Context they lack        4    you     what changed above them, what the
                                      reorg means, why the roadmap moved.

Running notes            -    shared  a shared doc, both editing, carried
                                      forward. Commitments with owners.
What does NOT belong: sprint status, ticket walkthroughs, code review.
Those have their own forums, and letting them in is how a one-to-one becomes
a status meeting the engineer starts dreading.

The ordering is the substance of the answer. Their agenda goes first because a one-to-one is their meeting, and if yours goes first the time runs out on your items every week and they learn that the slot is for you. The observable symptom of getting this wrong is an engineer who arrives with nothing prepared, which managers misread as disengagement when it is a rational response to a meeting they do not own.

Feedback belongs in every one, not in a quarterly cycle, because the cost of delivering it collapses once it is routine. If feedback appears only when something is wrong, its arrival is itself the message and the person stops hearing the content.

The reciprocal question — what should I do differently — is the part most managers skip and the part that produces the useful information. It rarely works the first four times. It starts working once you have visibly acted on something small, because you have then demonstrated that answering it is safe.

The running notes matter for an unglamorous reason: they are the evidence base for the promotion case and the performance conversation later. A manager who cannot produce dated specifics has a year of impressions instead, and impressions are where bias lives.

Where is the line between delegation and abdication?

Delegation transfers the work and keeps the accountability; abdication transfers both. The practical test is whether you can say, without asking, what the current state is, what the definition of done was agreed to be, and when you next expect to hear about it. If you cannot, you have not delegated, you have hoped. Delegation done properly names the outcome rather than the method, states the constraints that are not negotiable, says explicitly what decisions the person can make alone, and sets a checkpoint they own rather than one you impose. The failure mode on the other side is worth naming too: delegating the task while retaining every decision inside it, which is the most demoralising version because the person carries the effort and none of the authority.

How technical should an engineering manager stay?

Technical enough to judge, not to build. Concretely, that means you can read a design document and find the unasked question, follow a post-incident review without translation, tell whether an estimate is padded or fantasy, and hold an opinion about the team's architecture that survives contact with the people who own it. It does not mean owning production code on the critical path, because your calendar guarantees you will become the blocker. The direction of decay matters: credibility degrades slowly and invisibly, so managers who do nothing hands-on for two years usually think they are fine and are being managed around. Cheap ways to stay current are reviewing code without approving it, taking on-call shifts, and picking up unimportant work with no deadline attached to it.

Why is your first instinct to fix it yourself the wrong one?

Because it is nearly always faster in the moment and nearly always more expensive over the quarter. Taking the task back solves this instance, teaches the team that escalating to you resolves things, removes the learning that would have prevented the next instance, and adds to a queue that already has your name on everything that matters. The exception is real and worth stating so the answer is not absolutist: in an incident, or when a commitment fails this week, take the work. Then treat the fact that you had to as the actual defect and fix the cause — missing knowledge, an unclear owner, a system only one person understands. Doing that twice without follow-up is how a manager becomes the single point of failure they were hired to remove.

What is your output as a manager, and how would you measure it?

Your output is the output of your team plus the output of the teams you influence, which is Grove's formulation and remains the only definition that survives scrutiny. That is deliberately awkward, because it means nothing you did personally counts. Usable proxies are whether the team ships predictably enough that stakeholders stop asking, whether people are visibly more capable than a year ago and can name what changed, whether the team survives your absence for a fortnight, and whether attrition is voluntary in the direction you would choose. The measure to distrust is your own busyness. A manager whose calendar is full and whose team is blocked has high activity and no output, and the interview question behind this one is whether you can tell the two apart.

What does the first ninety days as a new manager look like?

Listen, then change one thing. The first fortnight is one-to-ones with everyone, asking what is working, what is broken, what they would fix if they could, and what they think my job is. Weeks three to six are for corroborating that against evidence — the incident log, the release history, the last two planning cycles — because the loudest account of a team is rarely the accurate one. Somewhere around week six, name one visible problem and fix it, because credibility comes from a delivered change rather than from a listening tour. What to avoid is a reorganisation in month one: you have inherited a set of arrangements that were rational for reasons nobody has told you yet, and undoing them early is how new managers spend trust they have not earned.

One-to-ones and feedback

What is a one-to-one for, and who owns the agenda?

It is for the things that will not be said in a group, and they own the agenda. That single ownership rule is what distinguishes it from a status update, and handing it over is not a courtesy — it is the mechanism, because the information you most need is the information they have to choose to give you. Your standing items are feedback, growth and context they lack from above, and they fit into the time that is left rather than the time you take first. Frequency beats depth: weekly for thirty minutes builds a channel that is open when something goes wrong, whereas monthly for an hour means the first fifteen minutes are spent re-establishing rapport and the hard thing never surfaces.

Show me feedback rewritten from vague to specific.

The same message three times, and only the third is usable.

VAGUE
  "You need to work on your communication."

  Unactionable. Communication with whom, doing what, and what would
  better look like. The person will either guess wrong or conclude that
  you find them generally deficient.

BETTER BUT STILL JUDGEMENT
  "You were quite dismissive in the design review."

  Names an occasion but describes character rather than behaviour, so it
  invites a defence of intent - "I wasn't being dismissive" - and the
  conversation becomes about whether you read them correctly.

SITUATION - BEHAVIOUR - IMPACT
  Situation  "In Tuesday's design review, when Priya proposed the
              queue-based approach..."
  Behaviour  "...you said 'we tried that, it doesn't work' and moved to
              the next agenda item without asking what was different
              about her version."
  Impact     "She didn't speak again in the meeting, and afterwards she
              asked me whether it was worth writing up the alternative.
              We may have lost an option because of thirty seconds."
  Ask        "What was going on for you there?"

The structure works because each part removes a specific escape route. The situation is dated and located, so the conversation cannot become an argument about whether this is a pattern. The behaviour is quotable and observable, so it cannot be denied — you are reporting what was said rather than diagnosing why. The impact is the part that motivates change, because most people are not trying to have that effect and did not know they had it.

The final question is not optional politeness. Roughly a third of the time the answer changes the feedback: they had context you lacked, they had already apologised, they were in the middle of something you did not know about. Skipping it turns a conversation into a verdict, and verdicts are argued with.

What to keep out: the word "always", any speculation about intent, and the sandwich. Wrapping criticism in two compliments teaches people that praise predicts bad news, so they stop hearing the praise and brace through it.

Timeliness is the other half. This feedback is worth giving on Tuesday afternoon and worthless in a review in March, because by then the specifics are gone and what remains is "your communication needs work" — the vague version you started with, arrived at by delay.

What makes feedback land?

Four things, and the absence of any one of them is enough to lose it. Specificity, so the person knows which behaviour on which occasion. Behaviour rather than character, because "you were dismissive" is arguable and "you moved past the proposal without asking a question" is not. Timeliness, since the value decays within days and the person's own memory of the occasion is the thing that makes it recognisable. And a relationship with enough deposits in it that criticism reads as investment rather than as a threat, which is why the manager who only appears when something is wrong cannot deliver feedback at all. The delivery detail worth adding is privacy for criticism and publicity for praise, reversed being one of the fastest ways to lose a team.

How do you take feedback from your own team?

By making it cheap to give and then visibly acting on something. Asking "any feedback for me?" at the end of a one-to-one reliably produces nothing, because the question is broad, the power gradient is real, and there is no evidence yet that answering it is safe. Narrower questions work: what did I get wrong in that decision, what should I have told you sooner, what am I spending time on that does not help you. Then the response matters more than the question — thank them, do not explain, and change something small within a fortnight so the causal link is visible. Defending yourself once closes the channel for a year, and the honest thing to say in an interview is that anonymous survey results are usually the first place you find out what people would not tell you directly.

What is the difference between coaching and directing?

Directing supplies the answer; coaching develops the person's ability to reach it. The choice is not a matter of style, it is a function of two variables: how urgent the outcome is and how much capability the person already has. High urgency or low capability means direct, clearly and without apology, because coaching someone through an outage or through a task they have never done is abandonment dressed as development. Low urgency and existing capability means coach — ask what they have considered, what they would do if you were on leave, what would have to be true for the other option to win. The failure in each direction is recognisable: managers who only direct build a team that cannot operate without them, and managers who only coach are experienced as evasive by people who wanted an answer.

Show me the same situation handled by coaching and by directing.

One scenario, two legitimate responses, and the point is what makes each correct.

Situation: a senior engineer proposes rewriting the reporting service in
a new framework. You think it is a poor use of a quarter.

DIRECTING                          COACHING
---------------------------------  ----------------------------------------
"We're not doing this quarter.     "Walk me through what problem the
The compliance deadline in         rewrite solves. What does it cost us
March needs both of you, and       to keep the current one for another
a rewrite puts it at risk. I'll    year? If we had six weeks rather than
own that decision. Let's revisit   a quarter, what's the version of this
in April - write down the case     that fits? And what would you cut from
now while it's fresh."             the roadmap to pay for it?"

when it is right                   when it is right
- the deadline is immovable and    - there is no immediate deadline
  the trade-off is not theirs      - they have the judgement and need
  to make                            practice at exercising it
- you hold information they        - the decision is genuinely theirs
  do not have                      - the real goal is that they own the
- ambiguity is doing harm            consequence either way

what it costs                      what it costs
- they learn less about the        - it is slower, and reads as evasive
  trade-off                          if you already know the answer
- repeated, it produces a team     - if you overrule them afterwards,
  that stops proposing              you have wasted their time and
                                     spent trust

The trap the table exposes is the hybrid: asking coaching questions when you have already decided. Engineers detect this immediately and it is worse than plain direction, because it costs them an hour and signals that consultation here is theatre. If the decision is made, say so and then explain the reasoning — the reasoning is the development, not the process of arriving at it.

The reverse error is directing on a decision that was theirs. Overriding a technical choice inside someone's own area, when nothing about the timeline forced it, transfers ownership of the outcome to you. If it then fails, they will not feel it, and you have removed the mechanism by which senior engineers become more senior.

The line worth saying in an interview is that you tell the person which mode you are in. "This one is mine and here is why" and "this is yours, I have opinions and they are not binding" are both fine, and the ambiguity between them is what does the damage.

Why is the annual review the wrong place for feedback?

Because everything useful about feedback has already expired by then. The specifics are gone, so what arrives is a summary the person cannot act on and cannot verify; the behaviour has had a year to become habit; and the stakes are now attached to money, so the person is negotiating rather than listening. The worst version is the surprise — a critical rating for something never raised in eleven months — which is a management failure being reported as an employee failure, and it reliably ends in either an exit or a formal dispute. The review should be a summary of conversations already had, containing no new information. If it contains news, the answer to give in an interview is that the cadence underneath it was broken and that is the thing to fix.

Performance and growth

What does a level on a career ladder actually mean?

Scope, not skill. Two engineers can be equally capable and sit at different levels because one is trusted with a component and the other with an ambiguous problem spanning three teams. Levels answer three questions: how large and how ill-defined a problem you can be handed, how much supervision the outcome needs, and how far your influence reaches beyond your own work. Framing it as scope is what makes the conversation tractable, because scope can be granted and observed whereas skill is asserted. It also produces the honest answer to the commonest grievance on any team — "I am as good as they are" — which is usually true and beside the point, since the question is what you have demonstrably owned rather than what you could do.

Show me two adjacent ladder levels as a table.

The gap between senior and staff is where most disputes happen, so it is the pair worth being able to state precisely.

                    Senior engineer (L5)          Staff engineer (L6)
------------------  ----------------------------  ---------------------------
Problem handed to   "Build the new pricing        "Pricing changes take three
you                 service to this spec"         weeks and cause incidents.
                                                  Fix that."

Ambiguity           Requirements are known.       Problem is a symptom. Part
                    You resolve technical         of the work is deciding what
                    unknowns.                     the problem actually is.

Blast radius        One service or component.     A domain across several
                                                  teams, or a cross-cutting
                                                  concern.

Supervision         Outcome reviewed. Approach    Trusted with the outcome.
                    is yours.                     You bring the framing to
                                                  your manager, not the plan.

Influence           Raises the bar on their       Changes how other teams
                    team. Mentors two or three.   work. Written artefacts
                                                  outlive the project.

Failure looks like  Late, or a design that        Solved the wrong problem
                    needed rework.                elegantly, and two teams
                                                  built on it.

Evidence            Shipped features, reviews,    Decisions others cite. A
                    incidents handled, people     migration completed. An
                    who improved near them.       argument settled durably.

The row that decides most promotion arguments is the first one. At senior level the problem arrives specified and the difficulty is in the solution; at staff level the problem arrives as a business complaint and defining it is most of the work. An engineer who is excellent at the second column's job but has only ever been given first-column problems is not being denied a promotion, they are being denied the scope that would produce the evidence — and that is the manager's failure to fix, not the engineer's to argue.

The influence row is the one candidates under-weight. Beyond senior, the ladder stops rewarding personal throughput almost entirely, and the fastest, most prolific engineer on the team can be genuinely stuck at senior for years. Saying that plainly, early, is kinder than implying that more output will eventually work.

Note that no row mentions years, headcount or technology. A ladder with "five years' experience" in it is measuring tenure, and a ladder that says "leads a team" has confused the individual-contributor track with management, which is the structural mistake that pushes good engineers into management to get paid.

The last row is the operationally important one. If a level's requirements cannot be evidenced from artefacts that already exist, the ladder cannot be applied fairly, and promotion decisions will be made on visibility instead — which is where bias enters a process that believes itself objective.

Show me a promotion case laid out as evidence.

A promotion case is not an argument that someone is good. It is evidence that they have already been operating at the next level.

Case: Priya R, Senior (L5) -> Staff (L6)
Period covered: Jan 2025 - Jun 2026. Calibration: 14 Jul 2026.

CLAIM 1  Operates on ambiguous, cross-team problems
  Given "pricing changes are slow and risky" with no spec. Produced the
  problem definition herself: 60% of the three-week cycle was manual
  spreadsheet reconciliation across Pricing and Billing, not code.
  Evidence  design doc PRC-14 (cited by two other teams' docs),
            cycle time 21 days -> 4 days, measured Mar and Jun.

CLAIM 2  Influence beyond her own team
  Wrote the rate-limiting standard now used by 9 services. Ran the
  migration as an opt-in with a paved path rather than a mandate.
  Evidence  adoption 9/11 services with no directive; ADR 022;
            two staff engineers on other teams asked her to review.

CLAIM 3  Trusted with the outcome
  Owned the Q1 pricing compliance deadline end to end, including the
  decision to descope the loyalty integration. Escalated once, with a
  recommendation attached rather than a question.
  Evidence  delivered 8 Mar, deadline 31 Mar. My involvement: two
            30-minute check-ins across the quarter.

CLAIM 4  Raises others
  Two engineers cite her as the reason they got unstuck on X. One was
  promoted to Senior in this period, with her review comments as the
  main development mechanism.
  Evidence  peer feedback from 4 named colleagues, quoted verbatim.

NOT YET DEMONSTRATED
  Has not had to hold an unpopular technical position against a senior
  stakeholder. Untested rather than absent - flagged so calibration
  weighs it honestly. Development plan: she leads the Q3 storage
  decision with me in the room, not speaking.

Four features make this a case rather than an endorsement. Every claim is a sentence from the ladder, so the panel is comparing evidence to a standard rather than to their impression of other candidates. Every claim has evidence that exists independently of the manager's advocacy — a document, a number, a named colleague — which means it survives the manager leaving. The period is stated, so nobody is promoted for one good quarter. And the case describes work already done at the next level, because promotion recognises demonstrated scope rather than granting it in advance.

The "not yet demonstrated" section is the part that gets cases through. A case with no gaps reads as advocacy and invites the panel to hunt for the omission themselves; naming it first makes the rest credible and turns the panel's objection into a development plan you have already written.

The claim to lead with is the one framed in business terms. "Twenty-one days to four" is arguable on its measurement and unarguable on its relevance, whereas "raised the quality bar" is unarguable and unassessable, and cases built entirely from the latter fail in calibration for reasons nobody can quite articulate.

What is absent is deliberate: no tenure, no comparison to peers, no mention of how much she wants it. All three are the standard ways an unfair promotion gets argued, and a manager who reaches for them in an interview has told you how their calibration works.

Why do promotion cases fail?

Usually because the evidence is of the current level performed excellently rather than of the next level performed at all, which is a failure of the manager's scoping over the preceding year rather than of the write-up. The other recurring causes are visibility — real impact nobody outside the team can attest to — and claims stated as qualities instead of artefacts, so a panel has nothing to weigh. There is also a timing failure worth naming: a case submitted at the first moment it is arguable, rather than at the point it is comfortable, spends the engineer's patience and the manager's credibility. The obligation this creates is to tell someone what is missing while there is still a year to fix it, because a first-time no in calibration is often followed by a resignation.

What is the difference between a growth plan and a performance plan?

Intent and consequence. A growth plan is developmental, aimed at the next level, owned by the engineer, and nothing bad happens if it slips. A performance plan concerns work that is currently below the expected standard for the level the person already holds, is owned jointly with a documented review date, and has a stated consequence if the standard is not met. Blurring them is a serious and common error: dressing a performance problem as development means the person believes they are on track until suddenly they are not, which is unfair and, in most jurisdictions, legally fragile. Say which one you are in, in the meeting, in those words. The person is entitled to know whether they are being developed or warned.

Show me a performance conversation scripted, including the opening line.

The first sentence decides how the next thirty minutes go, so it is worth having written rather than improvised.

Setting: booked 45 minutes, private, no agenda title that alarms. Written
summary sent within 24 hours. HR informed beforehand, not present.

OPEN - name the gap in the first sentence. Do not warm up.
  "I want to be direct, because I don't think I've been clear enough
   before now: your work over the last quarter has been below what we
   expect at senior level, and I'm worried about it. I want to spend
   this time being specific about what I mean and then hearing your
   view."

EVIDENCE - three dated examples, behaviour not character.
  "The payments migration slipped from March to May. Twice I found out
   we were behind from Dan rather than from you.
   The last two design reviews you attended, you hadn't read the
   document beforehand and the review got rescheduled.
   Three of your four PRs last month needed a second round of review
   for things the tests would have caught."

CHECK - then stop talking.
  "That's my view of the facts. Where am I wrong, and what am I
   missing about what's been going on?"

STANDARD - what good looks like, in observable terms.
  "At senior level I need to hear about a slip from you before it
   happens, designs reviewed with written comments, and PRs that land
   in one round most of the time."

PLAN AND DATE - support and consequence, both explicit.
  "Six weeks. Weekly 30 minutes on this specifically. I'll pair you
   with Anya on the next design. If we're not there by 8 September,
   the next conversation is a formal performance plan, and I'd rather
   have this one instead. I want you to succeed here, and I'm telling
   you now rather than in a review."

CLOSE
  "What do you need from me that you're not getting?"

The opening line does one job: it removes the possibility of the person leaving the room unsure whether this was serious. Managers soften the open because the conversation is unpleasant, and the reliable consequence is that the person hears a mild grumble, changes nothing, and is genuinely blindsided in six weeks — at which point the manager's own record shows they were told and the person's experience is that they were not.

The check step exists because you are often partly wrong. Half of these conversations surface something that reframes the evidence: an illness, a caring responsibility, an unstated dependency, or a piece of context that reveals the real problem is a badly scoped role. Finding that out after the plan is written is worse for everyone.

"I'd rather have this one instead" is the sentence that keeps the person working with you rather than against you. It states that the goal is recovery and not documentation, and it is only credible if the support named is specific and actually happens.

What to avoid: the sandwich, comparisons to named colleagues, anything about attitude, and any consequence you are not prepared to carry out. A deadline that passes with no action teaches the whole team what standards mean here, and the strong performers notice first.

Show me how someone is managed out with dignity.

If it reaches an exit, the quality of the process is judged by whether the person saw it coming and whether the team believes it was fair.

flowchart TD
    A[Concern observed<br/>written down with dates] --> B[Direct feedback in a one-to-one<br/>specific, behavioural, timely]
    B --> C[Expectations in writing<br/>standard, support and a review date]
    C --> D{Standard met by the date}
    D -->|yes| E[Return to normal management<br/>and say so explicitly]
    D -->|no| F[Formal plan with HR<br/>consequence stated in plain words]
    F --> G[Exit agreed with notice<br/>reference, dignity, no surprise]

The single most important property of this path is that no box surprises anyone. By the time the formal plan appears, the person has heard the same specific concern at least three times and has been told what happens if it does not change. Dignity is almost entirely a function of predictability — people can accept an outcome they saw approaching far better than a fair outcome that arrives from nowhere.

The yes branch is not decoration and is the branch managers handle worst. If the standard is met, say so unambiguously, in writing, and stop the process — because a manager who never closes a performance concern leaves someone working under permanent suspicion, and they will leave anyway, later, with the good will spent.

The exit itself is where dignity is actually delivered or lost. Full notice or pay in lieu, a factual reference, a story the person can tell that is true and not humiliating, and their choice about what the team is told. Time to hand over properly matters too, both because it is decent and because the alternative — a same-day escort — teaches everyone remaining exactly how much they are worth here.

The thing to say unprompted in an interview is that most of these cases are a hiring or scoping failure rather than a person failure, and that the honest review afterwards asks which. A team where three people have been managed out in a year does not have a talent problem.

What do you do with a strong engineer who is corroding the team?

Treat the behaviour as a performance problem, because it is one. The trap is weighing output against conduct as though they were separate ledgers, when the conduct is already reducing output — through the reviews people avoid, the ideas not raised, and eventually the resignations of the quieter engineers who leave without telling you why. So the conversation is the same as any other performance conversation, with specific incidents and a standard, and it says plainly that technical excellence does not buy an exemption. The part interviewers listen for is whether you would carry it through: tolerating this from a strong performer tells the whole team that the stated values are conditional on usefulness, and that lesson costs more than the person contributes.

Hiring and interviewing

Show me a hiring loop and what each stage actually assesses.

A loop is a sequence of independent signals, and the failure is stages that all measure the same thing.

flowchart TD
    A[CV and application screen<br/>is an hour worth spending] --> B[Hiring manager call<br/>role fit, motivation, constraints]
    B --> C[Technical screen<br/>can they build working software]
    C --> D[Design or domain deep dive<br/>judgement under ambiguity]
    D --> E[Behavioural interview<br/>evidence of how they have worked]
    E --> F[Debrief with scores written first<br/>hire or no hire]
    F --> G[Offer and close<br/>owned by the hiring manager]

The design principle is one signal per stage, assessed against a written rubric, and no stage repeating another. Loops decay in a predictable direction: every interviewer gravitates towards coding, because it is the easiest thing to run and the most comfortable to defend, and after a year you have four coding rounds and no evidence about collaboration, ownership or how the candidate behaves when they are wrong.

The behavioural stage is the one treated as filler and it is the one that predicts tenure. Run it as structured past-behaviour questions — tell me about a time you disagreed with a technical decision and lost — with follow-ups that push for specifics, because a rehearsed narrative falls apart at the third "what did you actually say". Hypotheticals belong in the design round, not here: what someone would do is a measure of their taste, and what they did is a measure of their behaviour.

Scores written before the debrief is the detail candidates omit and the one that does the most work. In an unstructured debrief the first confident speaker sets the group's position and everyone converges, so independent written scores are the mechanism that preserves the disagreement worth having.

The last stage belongs to the hiring manager rather than a recruiter. A candidate deciding between offers is deciding about the person they will work for, and delegating the close reliably loses candidates for reasons that are recorded as compensation and are not.

Show me an interview scorecard structure.

The purpose of a scorecard is to force the evidence to be written down before the conclusion is reached.

Candidate: A. Okafor      Role: Senior BE       Stage: design deep dive
Interviewer: G. Sharma    Date: 21 Jul 2026     Submitted BEFORE debrief

Competency assessed in this stage - three, no more.

1  Decomposes an ambiguous problem
   Rating   4 / 4 = strong yes, 3 = yes, 2 = no, 1 = strong no
   Evidence "Asked about read/write ratio and retention before drawing
             anything. Named two hard requirements and put notifications
             explicitly out of scope."
   Against  L5 rubric line: 'resolves technical ambiguity independently'

2  Names the cost of their own choices
   Rating   2
   Evidence "Added a cache and a queue without prompting. When asked how
             the cache is invalidated, said 'TTL' and did not mention
             stale reads until I pushed twice. Did not raise consumer lag
             at all."
   Against  L5 rubric line: 'anticipates failure modes of their design'

3  Responds to new information
   Rating   3
   Evidence "When I added the multi-region requirement, revisited the
             earlier storage choice out loud rather than bolting on.
             Said 'that changes the answer I gave you before'."

Overall  HIRE at L5, not L6. Reservation: operational depth.
Specific follow-up for the next interviewer: ask how they knew a system
they built was failing in production, and who was paged.

NOT ASSESSED HERE: coding, collaboration, domain knowledge.

Three features make this work. Ratings are on an even scale with no middle, so "3" cannot mean "I would rather not commit" — the commonest way a loop produces no decision. Evidence is quoted rather than characterised, so the debrief can disagree with the interpretation while sharing the facts. And each competency cites the rubric line it maps to, which is what stops the bar drifting with the market and the mood.

The line naming what was not assessed is the one candidates never mention. Without it, a debrief silently treats absence of signal as negative signal, and the strongest engineer in the pool is rejected because nobody asked about the thing they are best at.

The follow-up note is how a loop compounds rather than repeating itself. Passing a specific unresolved doubt forward means the last interviewer resolves it instead of re-running round one, which is also how a loop gets shorter without getting worse.

The "before debrief" stamp is a process control, not bureaucracy. A scorecard written afterwards is a rationalisation of a decision already made in the room, and it will read as evidence to everyone who sees it later, including a tribunal.

What actually reduces bias in hiring?

Structure, and almost nothing else that is commonly proposed. The interventions with real support are the same questions asked of every candidate in the same order, a rubric written before the candidate was seen, evidence recorded before the recommendation, independent scores collected before any discussion, and a work-sample task resembling the actual job. Unconscious-bias training on its own has a poor record and mainly produces confidence. What matters more than any of these is where the pipeline comes from, since a loop cannot correct for a candidate set drawn entirely from one network. The mechanism behind all of it is the same: bias flourishes wherever judgement is unconstrained and undocumented, so you reduce it by reducing the number of moments that describes.

How do you calibrate an interview bar across interviewers?

By making the standard visible and the disagreements explicit. Practically: a written rubric with example answers at each rating, shadowing before anyone interviews alone, reverse-shadowing before they are signed off, and periodic sessions where several interviewers score the same recorded or written response and then discuss the gaps. Then watch the data — an interviewer whose pass rate is double or half the team's is either miscalibrated or asking a different question than they think. The failure to name is the "high bar" self-image: an interviewer who rejects nearly everyone is often read as rigorous, when they may simply be asking a question with one acceptable answer. Both drift directions cost, and only one of them is ever noticed.

How do you decide on a borderline candidate?

You do not hire them, and the reasoning has to be better than an instinct. The asymmetry is real: a mediocre hire costs a year of management attention, drags the bar, and is far harder to reverse than a slow requisition. But "no" for the wrong reasons is also expensive, so first check what is actually borderline. If it is a gap in one skill that the team already has and can teach, that is a development plan rather than a rejection. If it is doubt about ownership, judgement, or how they behave when disagreed with, that is a no, because those change slowly and under nobody's control but their own. And if the loop simply did not gather the signal, the answer is one more targeted conversation rather than a coin toss.

What does onboarding owe a new hire in the first thirty days?

A shipped change in week one, a named person to ask stupid questions, and clarity about what success looks like at ninety days. The first is the load-bearing item: a small change reaching production in the first few days proves the environment works, teaches the release path, and converts a passive week of reading into membership. Beyond that, they need the map that is not written down — who decides what, which systems are dangerous, which documents are lies — which is why a buddy outperforms a wiki. The measurable failure is the two-week induction with no commit, after which people quietly conclude the team is disorganised, and that impression is remarkably durable. Onboarding is also the cheapest retention work available, since the strongest predictor of an early exit is a bad first month.

What is a structured interview, and why does it beat a conversation?

A structured interview asks predetermined questions in a fixed order and scores answers against a rubric defined beforehand. It beats an unstructured conversation on predictive validity by a wide margin in the research, and the reason is mundane: a conversation measures similarity to the interviewer, and it feels like insight because rapport is pleasant. Structure also makes a loop improvable, since you can look at which questions distinguished good hires from bad and change them, whereas you cannot audit a chat. The objection that it feels robotic is worth answering directly: the questions are fixed and the follow-ups are not, so the conversation is still a conversation — it simply starts from the same place for everyone.

Team health and delivery

Show me a team health check as signals with thresholds.

A health check is only useful if each signal has a number at which you act, otherwise it is a mood board.

signal                     healthy      watch        act now
-------------------------  -----------  -----------  --------------------------
Voluntary attrition        <10%/yr      10-15%       >15%, or 2 in a quarter
                                                     from one sub-team

Time to first commit       <5 days      5-10 days    >10 days for a new joiner

Change failure rate        <15%         15-25%       >25%, or any upward trend
                                                     over 3 months

Unplanned work             <20%         20-35%       >35% of sprint capacity

On-call pages out of       <2/week      2-5/week     >5, or the same alert
hours                                                twice in a month

PR review wait, p75        <4 hrs       4-24 hrs     >24 hrs, or one person
                                                     on >40% of reviews

Meeting-free blocks        >3/wk of     1-2 blocks   0 blocks of 3+ hours
per engineer               3+ hours

One-to-ones held           >90%         80-90%       <80%, or cancelled by
as scheduled                                         the manager

Engineers who can name     100%         -            anyone who cannot
their next growth step

The pattern across the table is that the leading indicators are all about time and interruption, and the lagging one is attrition. By the time attrition moves, the cause was six months ago and is usually visible in the review-wait row or the out-of-hours pages, which is why acting on the cheap rows is the whole point of keeping the list.

Thresholds have to be defined before the number is bad. A threshold agreed while you are already over it becomes a negotiation about whether the threshold was right, and that conversation always concludes that the current state is fine.

The two rows that look softest carry the most information. A single reviewer handling forty per cent of pull requests is both a bottleneck and a bus-factor risk, and it is invisible in every delivery metric until that person takes leave. One-to-ones cancelled by the manager is the row that predicts the others, because it is what a manager sacrifices first and it is the mechanism by which they would otherwise have learned that something was wrong.

The honest caveat: this is a dashboard for the manager, not a scorecard for the team, and the moment it is reported upward as a rating the numbers start being managed rather than the team. Say that before an interviewer asks it.

Show me the DORA metrics and how each gets gamed.

Four metrics, each individually gameable, which is why they are quoted as a set.

metric                what it measures        how it gets gamed
--------------------  ----------------------  ---------------------------
Deployment            releases to production  deploy trivial or empty
frequency             per unit of time        changes; split one change
                                              into six PRs; count deploys
                                              to staging

Lead time for         commit to running in    start the clock at merge
change                production               rather than at first commit,
                                              hiding two weeks of review
                                              and queueing

Change failure        % of deploys causing    reclassify incidents as
rate                  degradation or rollback "planned maintenance"; fix
                                              forward so no rollback is
                                              recorded; raise the bar for
                                              what counts as an incident

Time to restore       incident start to       start the clock at
service               service recovered       acknowledgement rather than
                                              at customer impact; close
                                              the incident at mitigation
                                              and track the real fix
                                              separately

The pairing that resists gaming:
  throughput  = frequency + lead time
  stability   = failure rate + time to restore
  Improving either pair alone is nearly always a trade, so the honest
  claim is movement in one pair without regression in the other.

Every one of these is gamed by moving a definition rather than by lying, which is what makes it hard to catch and easy to do without intending to. The defence is that the definitions are written down, owned outside the team being measured, and instrumented from systems rather than self-reported — deployment records and incident timestamps, not a spreadsheet somebody fills in.

The structural point is that these are diagnostic, not a target. Goodhart applies immediately: attach a bonus or a promotion to deployment frequency and you will get deployment frequency, delivered by exactly the mechanisms in the right-hand column, at the cost of the thing the metric was proxying for. Used properly they tell a team where its own bottleneck is, which is why the useful question is not "what is our lead time" but "which part of it is queueing".

What DORA does not measure is worth volunteering, because it is where the interesting management questions live. It says nothing about whether the right thing was built, nothing about team health or sustainability, and nothing about the quality of the design — a team can be elite on all four while shipping features nobody uses, and that failure is more expensive than a slow pipeline.

The fifth metric added later, reliability, is worth knowing by name, and so is the context: the benchmarks come from a survey of self-reported data, so comparing your team to the elite cluster is a much weaker exercise than comparing your team to itself last quarter.

Why do people really leave?

Rarely for money, and almost never for the reason on the exit form, because the exit form is filled in by someone who wants a reference. The recurring causes are a manager they do not trust or learn from, work with no visible growth or arc, a promotion that was implied and did not arrive, being on a team that ships nothing they are proud of, and sustained overload — particularly on-call — that nobody acted on when they raised it. Money is usually the trigger rather than the cause: an external offer resolves a decision already made months earlier. The practical implication for a manager is that a resignation is a lagging indicator, and the leading ones are available for free in a one-to-one to anyone who asks the question and does not defend the answer.

Show me a delivery forecast expressed as a range with confidence.

A single date communicates certainty you do not have, so the forecast is a distribution and the conversation is about which end of it to plan against.

Feature: multi-currency checkout
Remaining scope: 34 stories, decomposed to <=3 days each

Throughput, last 8 sprints: 11, 9, 14, 8, 12, 10, 7, 13
  median 10.5/sprint   worst 7   best 14

Naive plan     34 / 10.5 = 3.2 sprints -> "6 weeks, so 8 September"
                                          <- this is the answer that
                                             later becomes a slip

Monte Carlo, 10,000 runs resampling those 8 sprints with replacement.
Read each row as "chance of being finished by", which is what the
simulation produces. Do not pick a confidence first and read a date
off it - the model does not owe you a 50/75/85/95 ladder.

  finished by          34 stories   39 stories (scope +15%)
  ------------------   ----------   ----------------------
  3 sprints,  8 Sep        32%               4%
  4 sprints, 22 Sep        97%              77%
  5 sprints,  6 Oct       >99%              99.8%

The curve is steep because this team's throughput is steady: eight
sprints between 7 and 14, so three sprints is almost never enough and
four almost always is. There is no meaningful 85% date between them.

Assumptions, stated because they are the actual risk:
  - scope does not grow. Historically it has grown ~15% mid-flight,
    which is the second column: 22 Sep drops from 97% to 77%.
  - two engineers, no holiday. Anya is out for 2 weeks in September:
    subtract ~1 sprint of throughput from the sampled range.
  - the payment provider's sandbox is available by 11 Aug. This is
    someone else's dependency and the single largest risk in the model.

What I would say out loud: "22 September, 97% if the scope holds and
77% if it grows the way it usually does. 8 September is a one-in-three
shot and I would not plan a marketing launch on it. If you need a date
that survives scope growth, it is 6 October. The provider sandbox is
the thing that could move all of this by a month, and I need a date
from them by Friday."

The essential move is that the forecast is built from measured throughput rather than from estimates of effort. The team's last eight sprints already contain everything estimation tries to guess at — interruptions, holiday, review latency, the work nobody planned — so sampling from history is both more honest and less effort than re-estimating.

The shape of the curve is the information the stakeholder actually needs, and here the shape is steep rather than wide. Two weeks separates a one-in-three chance from a near-certainty, because eight sprints of throughput between seven and fourteen do not admit much doubt about whether four sprints is enough. Say that plainly rather than manufacturing intermediate dates: a table that offers a 75% date and an 85% date one row apart, both landing on 22 September, is decorating a two-point distribution and a numerate stakeholder will notice.

Where the real width comes from is the second column. Scope growth moves 22 September from 97% to 77% — a far bigger effect than throughput variance — which is the honest answer to what could go wrong. Presenting one number hides both effects and guarantees that the conversation happens later, during a slip, when your credibility is the thing being spent.

The assumptions block is where the forecast earns trust. Two of the three are outside the team's control, and naming the provider dependency in advance changes what the stakeholder does — they can chase it — whereas discovering it in September only changes who is blamed.

The pushback to rehearse is "just give me a date". The answer is to give one, with its confidence attached, and to say what you would do differently at each level: at one-in-three I would not book the launch, at ninety-seven I would. That converts a demand for false certainty into a decision about risk appetite, which is the stakeholder's decision to make and not yours to absorb silently.

Why is a single-date commitment dishonest?

Because delivery is a distribution and a date is a point, so quoting one without a confidence level asserts a certainty that no model of the work supports. The mechanics are worse than they look: the point usually quoted is the median or better, which means a coin-flip chance of missing it, and everyone downstream plans as though it were a guarantee. What follows is predictable — the slip arrives late, because the team was optimistic until the last fortnight, and the lesson learned upward is that engineering cannot be trusted rather than that the question was badly formed. The alternative is not refusing to commit. It is committing at a stated confidence, naming the assumptions that would move it, and reforecasting on a fixed cadence so a slip is reported early and small.

How do you protect focus time?

By treating calendar fragmentation as a delivery problem with an owner rather than as each engineer's personal discipline problem. The moves that work are structural: no-meeting blocks that are defended by the manager, a default of asynchronous written updates in place of status meetings, a rota so one person handles interrupts and support for the week instead of everyone being interruptible, and a standing willingness to be the person who declines a meeting on the team's behalf. The reason it is the manager's job is asymmetry of cost — an engineer declining a director's invitation pays a price you do not — so delegating the protection to them means it will not happen. The measure is blocks of three hours or more per engineer per week, because two one-hour gaps are not the same resource.

What makes an on-call rota fair?

Predictability, compensation and a real path to reducing the load. Predictability means the schedule is published far enough ahead to plan a life around, swaps are easy, and nobody is on call alone on their first rotation. Compensation means paid or time off in lieu, applied consistently, because unpaid out-of-hours work is a tax levied on whoever has the fewest domestic commitments. The path to reduction is the part that decides whether the rota is sustainable: pages that fire without requiring action get deleted or fixed within the sprint, and time to do that appears on the board rather than in goodwill. The signal that it has gone wrong is the same senior engineer being paged for the same alert quarterly, which is a capacity decision the organisation is making without admitting it.

How can you tell whether an incident culture is genuinely blameless?

Look at behaviour rather than at the stated policy. In a genuinely blameless culture people volunteer their own mistakes in the review, junior engineers declare incidents without checking upward first, the write-up names systems and missing guardrails rather than individuals, and remediation items have owners and dates and actually get done. The tells for the opposite are a review that concludes with "engineer did not follow the runbook", a reluctance to page anyone senior, and action items that are all variations on being more careful. The sharpest question to ask a team is when someone last said "I caused this" out loud in a review, because in a blaming culture the answer is never, and what you lose is not comfort but the information you needed to prevent the recurrence.

Working with other functions

Show me a priority trade-off presented to a stakeholder as options.

The move is to stop arguing about whether something is possible and present the choices with their costs attached.

Ask: "can we add SSO for the enterprise deal, in Q3?"
Team capacity in Q3: roughly 9 engineer-weeks after on-call, holiday
and the committed compliance work. SSO is estimated at 6-8.

OPTION A  Take SSO, delay the reporting rebuild to Q4
  cost      reporting slips one quarter. Three named customers have
            asked for it; two are in renewal in November.
  gains     enterprise deal unblocked, ~£240k ARR per sales
  risk      if SSO runs to 8 weeks there is no slack for the
            compliance work, which has a fixed March date

OPTION B  Take SSO, drop the two smaller roadmap items
  cost      in-app notifications and the CSV export go to the backlog
            and realistically do not return this year
  gains     same deal, compliance work untouched
  risk      lowest of the four. My recommendation.

OPTION C  Do not take SSO this quarter
  cost      deal likely lost or delayed to Q1
  gains     roadmap intact, team not switching context
  risk      none to delivery. All to revenue.

OPTION D  Take SSO with a contractor
  cost      6 weeks' contract, plus ~1.5 weeks of our senior engineer's
            time to onboard and review. Net saving is smaller than it
            looks, and we own the code afterwards.
  gains     roadmap mostly intact
  risk      auth is a poor first task for someone who does not know
            the system. I would not choose this for security-adjacent
            work.

What I need from you: a decision between A and B by Friday, because
the reporting team starts branching on Monday. I am not asking you to
decide how, only which of these costs you would rather pay.

The structural point is that none of the options is "no". Refusal invites escalation and positions engineering as the obstacle, whereas four priced options move the decision to the person who owns the trade-off between revenue and roadmap — which is genuinely not the engineering manager's decision to make.

The recommendation is named. Presenting options without one is a common over-correction that reads as abdication: you have more information than the stakeholder about relative risk, and withholding your view to appear neutral makes the decision worse. Say which one and why, then accept the answer.

Option D is included because someone will suggest it, and pricing it honestly — including the review overhead and the fact that auth is the wrong first task — is more persuasive than dismissing it when it is raised in the meeting.

The last paragraph does the work that most of these conversations lack: a named decision, a named decider and a date, tied to a real consequence of delay. Without it the options document circulates, nobody decides, and the default outcome happens by drift — which is nearly always the worst of the four.

How do you say no to a stakeholder?

By pricing it rather than refusing it, because "no" invites escalation and a price invites a decision. So the answer is "yes, and here is what it displaces" with the displacement named specifically enough to be traded: this feature costs six weeks and the reporting rebuild is what moves. Then ask what problem the request was solving, since requests arrive as solutions surprisingly often and there is frequently a cheaper approximation that satisfies the actual need. Two conditions make this credible. Your estimates have to have been roughly right before, or the price is read as resistance. And you have to actually say yes sometimes, because a manager whose every answer is a cost is routed around within two quarters.

What does a product manager owe you, and what do you owe them?

They owe you a prioritised problem rather than a specification, the reasoning behind the priority, early visibility of what is coming so the technical work can be sequenced, and a decision when you need one rather than a deferral. You owe them honest forecasts with their uncertainty attached, early warning of a slip rather than a heroic recovery attempt, options with costs instead of a flat no, and the technical constraints translated into consequences they can weigh. The relationship fails in two symmetrical ways: a product manager who specifies solutions removes the engineering judgement they hired, and an engineering manager who hides risk until it is unrecoverable removes the product manager's ability to manage expectations. The functional version disagrees openly and presents one position outward.

Show me an escalation path.

An escalation path exists so that raising something is a defined act rather than a judgement call about whether you are being annoying.

flowchart TD
    A[Engineer is blocked] --> B[Team channel<br/>15 minutes, anyone may unblock]
    B --> C[Tech lead or manager<br/>same day, decision or reprioritise]
    C --> D[Peer manager of the blocking team<br/>within one working day]
    D --> E[Shared manager above both<br/>only with a written ask attached]
    C --> F[Incident channel<br/>if customers are affected right now]
    E --> G[Director forum<br/>a decision nobody below could make]

The purpose of writing it down is to remove the social cost. Most escalation failures are not disputes about the right level, they are an engineer sitting on a blocker for three days because raising it felt like complaining, and a documented path with times attached converts that into following the process.

Each step has a time bound, and the bound is the actual mechanism. "Escalate if this is not resolved by tomorrow" is a rule; "escalate if it feels stuck" is a personality test that extroverts pass.

The written ask at the fourth step is what stops escalation becoming complaint. By the time something reaches a shared manager it should arrive as a specific request — we need two weeks of the platform team in August, here are the two options and what each costs — because a manager handed a grievance can only convene a meeting, while a manager handed a decision can make it.

The branch to the incident channel matters because customer impact is not an escalation, it is a different process with a different clock. Conflating the two is how an outage spends forty minutes going up a management chain.

The last thing to say is that a path used constantly indicates a structural problem rather than a healthy culture. Repeated escalation to the same boundary means the dependency between those two teams is misplaced, and the durable fix is in the org chart or the service boundary rather than in faster escalation.

How do you communicate a slip upward?

Early, in writing, with a revised forecast and the decision you need. The structure that works is four sentences: what changed, what the new range is, what the options are, and what you are asking for. What loses trust is not the slip, it is the timing — a date missed the week it was due, after weeks of green status, tells your manager that your reporting is worthless and every future estimate gets discounted. So the discipline is to report the risk when the probability moves, not when the outcome is certain, and to be visibly wrong sometimes in the optimistic direction. Never bring a slip without options, and never bring one with an implied request for more people, since that is the response most likely to make the date worse.

How do you disagree with your own manager?

In private, on the consequence rather than the choice, with the information they may not have. "This commits us to a migration we cannot reverse, and here is what it costs in Q4" gives them something new, whereas "I think that is the wrong call" offers a competing opinion and invites a contest of seniority. Ask what they know that you do not, because senior managers routinely hold commercial or political context that makes an apparently poor decision reasonable. Then, if it goes against you, disagree and commit — visibly, because a manager who relitigates a lost decision in front of their team damages more than the decision did. Record your reasoning in writing so the position is recoverable if the risk materialises, and do not say you told them so if it does.

What do you owe your skip-level that you do not owe your team?

An honest view of risk and of individual performance, unsoftened. Your manager's manager needs to know which commitments you are genuinely unsure of, which people are struggling, and where your team is a single point of failure — information you cannot broadcast downward without doing harm. The reverse also holds: decisions still being debated above, other people's compensation, and anything told to you in confidence are not yours to pass down, and a manager who leaks upward turbulence to look transparent destabilises the team for their own comfort. The line worth articulating is that you filter but do not distort. Withholding unformed information is judgement; telling the team the roadmap is safe when you know a reorg is coming is a lie you will be present for the discovery of.

Difficult situations

How do you manage a former peer?

By naming the change explicitly and early, in a conversation rather than by behaving differently and hoping it is noticed. Say what is different — that you now decide some things you used to argue about, that you have information you cannot share, that their compensation and progression are partly yours — and ask what they want from you. The two failure modes are opposite and both common: over-correcting into formality, which reads as betrayal to someone who was a friend on Friday, and failing to correct at all, so the first hard decision arrives as a shock and the rest of the team watches you avoid it. The test the team applies is whether your former closest colleague gets the easiest reviews and the best work, and they will have decided within a month.

How do you manage someone more senior technically than you?

By being clear that you are not there to out-argue them, and then being useful. Your value is scope, context, unblocking and career leverage rather than technical authority, and stating that plainly at the start removes the contest they may be expecting. Delegate the technical decision genuinely — including the ones you would have made differently — and hold them to the same accountability as anyone else on outcomes, communication and how they treat the team. Where you do disagree technically, ask questions rather than asserting, and be willing to lose in public. What destroys the relationship is pretending to a judgement you do not have, which they detect immediately. The honest framing for an interview is that a manager who needs to be the best engineer in the room cannot manage staff engineers at all.

How do you handle a reorg you did not choose?

Communicate quickly, own the message, and separate what is decided from what is not. Silence is the most expensive option because the team fills it with worse versions than the truth, so say what you know, say what you do not know, and say when you will next say something — then keep that date even if the answer is still nothing. Do not distance yourself from the decision to preserve your own standing: "they have decided" invites the team to conclude their manager has no influence, which is more destabilising than an unpopular change. Expect performance to dip for a quarter, protect the one or two people most likely to leave with a direct conversation about what they want next, and re-establish the ordinary rhythm early because predictability is the thing people have just lost.

How do you handle the departure of a key person?

Treat the knowledge concentration as the incident and the resignation as the symptom. In the notice period, prioritise transfer over delivery — documented runbooks, paired work on the systems only they touch, a written list of what they know that nobody else does, produced by them — and accept the sprint cost, because the alternative is paying it repeatedly for a year. Resist the instinct to backfill identically and immediately; the departure is the cheapest opportunity you will get to reconsider the shape of the team. Handle the exit generously, since how you treat someone leaving is watched closely by everyone staying and determines whether they return. Then ask the uncomfortable question of why they left, and whether you knew, because if it surprised you the one-to-ones were not working.

What do you do about two senior engineers who cannot work together?

Establish what the disagreement actually is, because it is usually a technical dispute with no owner that has been left to fester into a personal one. Talk to each separately, then together with a specific agenda rather than an invitation to air grievances. If the root is a genuine technical disagreement, the fix is a decision — made by you if necessary, written down with the reasoning, and closed so it stops being relitigated weekly. If the root is conduct, it becomes a performance conversation with each of them, individually and specifically. What does not work is waiting, or restructuring the work so they never interact, which resolves the symptom and teaches the team that avoidance is rewarded. If it persists after a clear expectation, one of them moves.

What do you do when you inherit a team that is failing?

Find out what kind of failing it is before doing anything, because the four common causes need opposite responses. If the team cannot ship, the cause is usually process or a dependency and the fix is external. If it ships the wrong things, the problem is upstream in product and the fix is a relationship you do not yet have. If it is demoralised, the cause is often a previous manager and the fix is slow and mostly consists of keeping small promises. If capability is genuinely absent, that is hiring and performance work measured in quarters. Then pick one visible problem and fix it within about six weeks, because a failing team's first question is whether you can actually change anything, and no amount of listening answers it.

How do you handle a quietly disengaged engineer?

Name what you have observed and then be quiet. "You have been quieter in reviews and you have not picked up anything beyond your tickets for a month — what is going on?" is a question that can be answered, whereas "is everything okay?" reliably produces "fine". Then listen for which of the usual causes it is: a promotion that did not arrive, work they find meaningless, something outside work entirely, or a decision to leave that has already been made. The responses are genuinely different, and guessing wrong is worse than asking. The thing to avoid is treating disengagement as a performance problem in the first conversation, because it usually is not one yet, and framing it that way converts a recoverable situation into a resignation.

Interview traps

Show me a situation-task-action-result answer for a failure question.

Failure questions are answered badly because candidates optimise for looking good, which produces a story with no failure in it.

Q: "Tell me about a time you failed as a manager."

SITUATION
  "I took over a team of seven in 2024 that owned billing. Two of the
   seven were the only people who understood the invoicing engine, and
   both had been there four years."

TASK
  "My priority was a migration off a deprecated tax provider with a
   hard March deadline. Retention was not on my list."

ACTION - what I actually did, including the wrong part
  "I put both of them on the migration because they were the only ones
   who could do it, and I protected the deadline by shielding them from
   everything else. In our one-to-ones I asked about the migration. When
   one of them said he was 'a bit fed up with billing', I heard it as
   normal grumbling about a hard quarter and said the migration would be
   over in six weeks. I did not ask what he wanted next, and I never
   revisited it."

RESULT - the cost, stated plainly
  "He resigned three weeks before the deadline. We hit the date, but the
   next two quarters cost us: the remaining engineer was effectively on
   call for invoicing alone, we had two customer-visible billing
   incidents in April and May, and the backfill took five months to
   become productive. I also had a second person tell me later that she
   had been considering leaving for similar reasons and I had no idea."

WHAT I CHANGED - specific and testable
  "Two things. Every one-to-one now has a standing question about what
   they want next, asked as a real question rather than annually - which
   is how I found out about the second person. And I track knowledge
   concentration explicitly: any system with fewer than two confident
   owners is a risk on my board with a name against it, and I now spend
   throughput on that deliberately rather than treating it as slack.
   I would still have staffed the migration that way. What I would not
   do again is hear 'a bit fed up' and move on."

The answer works because the failure is real and the cost is quantified. A resignation, two incidents and five months of ramp are consequences an interviewer can weigh, whereas "it was a learning experience" is a claim about the candidate's attitude rather than about anything that happened.

The action section contains the actual mistake, in the first person, without an organisational culprit. This is the section candidates hollow out — the story becomes something the company did to them — and the hollowing is obvious from outside, because a failure with no decision of yours in it is not your failure.

The distinction in the final paragraph is the senior move. Separating the decision that was defensible with the information available from the specific signal that was missed shows judgement rather than blanket contrition, and it produces a lesson narrow enough to have changed behaviour. "I would still staff it that way" is what makes the rest believable.

The change described is verifiable. A standing question and a risk register entry are things an interviewer can probe — what did it surface, show me the current list — which is exactly why they are worth stating, and why "I communicate more now" is not.

What is an interviewer testing when they ask about a time you failed as a manager?

Whether you can hold accountability without either defending yourself or performing contrition, and whether the lesson was specific enough to have changed what you do. A strong answer names a real decision of yours, states what you knew at the time and why the decision was defensible then, identifies the signal you missed, quantifies the cost as it actually landed, and describes a concrete change — a question now asked routinely, a threshold now tracked. A weak answer picks something safe, blames a reorg or a stakeholder, or offers a failure that is secretly a boast about ambition. The secondary thing being assessed is whether you can tell an unflattering story about yourself calmly, because a manager who cannot do that in an interview cannot do it in an incident review either.

Why is "I hire great people and get out of the way" a weak answer?

Because it describes an absence of management and claims it as a philosophy. Every concrete thing a manager does — setting context, deciding priority, giving feedback, growing people, protecting focus, handling the person who is struggling — is precisely the getting in the way that the phrase disclaims, so the answer tells an interviewer either that you do not do those things or that you cannot articulate them. It also fails the obvious follow-up: what happens when a great person underperforms, or when two of them disagree, or when the priority is wrong? The defensible version of the same instinct is specific about autonomy — I own the what and the why, they own the how, and here is where I intervene anyway — because that names a boundary rather than avoiding one.

What do interviewers hear when every success is described as yours?

That they are about to hire someone whose team will not stay. Management answers given entirely in the first person — I delivered, I decided, I fixed — read as either an inaccurate account or an accurate one about someone who takes credit, and both are disqualifying for a role whose output is other people's work. The opposite over-correction is nearly as bad: an answer entirely in "we" leaves the interviewer unable to tell what you personally did, and "the team decided" invites the suspicion that you were present rather than accountable. The pattern that works is first person for your decisions and third person for the outcomes: I scoped it this way, I made this call and it was wrong, they built it and Priya found the flaw that saved us.

How should you answer "how do you handle a low performer"?

With a process and a specific case, in that order, and with the word "or" in it somewhere. The process is: establish whether the standard was ever made explicit, check whether the problem is capability, clarity, motivation or fit, give direct dated feedback, put expectations and support in writing with a review date, and then either return to normal management or move to a formal plan. The case is what proves you have done it. Two things interviewers listen for specifically: whether you consider that the role or the scoping may be the defect rather than the person, and whether any of your examples ended in recovery — because a manager whose every story ends in an exit is describing a filter, not management, and one whose stories never end is describing avoidance.

Why is "it depends on the person" a weak answer to a management question?

Because it is true, unfalsifiable and declines the question. It tells the interviewer nothing about your judgement, since every management answer depends on the person, and it is the most common way candidates avoid committing to a position they might have to defend. The strong version names what it depends on, gives the answer for each branch, and then commits: this depends on whether they lack the skill or the motivation — if it is skill I pair them and set a checkpoint, if it is motivation the conversation is about what they want and whether this role can offer it, and in the case I described it was the second. The underlying pattern is the same as in a design round: name the variable, resolve it with a stated assumption, then decide.

What single question most reliably separates candidates in an engineering management round?

"Tell me about someone you developed, and what specifically you did." It resists preparation because the answer has to be about a named individual over a real period, and it exposes the difference between a manager who has grown people and one who has supervised them. A strong answer states where the person started, what the specific gap was, what the manager did that a passive manager would not have done — the scope deliberately granted, the uncomfortable feedback given, the project handed over at the cost of it going slower, the promotion case built and argued — and what the person can do now that they could not before. It is comfortable naming the part that failed or took longer than expected. A weak answer describes someone who was already good and got promoted, with the manager present but not causal, or it retreats into process: I ran weekly one-to-ones and we discussed their growth plan. The underlying test is whether developing people is something the candidate does deliberately, at a cost they can name, or something they believe happens near them.