How would you design the interview loop for a senior engineer on your team?
Work backwards from the four or five competencies the role genuinely needs, map each to exactly one round so nothing is measured four times, score against shared behavioural anchors, require written feedback before the debrief opens, and treat both false negatives and candidate experience as data you owe yourself.
What the interviewer is scoring
- Does the candidate derive the rounds from a competency list, or assemble a loop from the interviews their team already likes running
- Whether the loop has a coverage matrix with no competency measured twice by accident and none left uncovered
- That they can explain what a rubric anchor looks like, rather than defending a bare one-to-four scale
- Whether independent written feedback before the debrief is presented as a mechanism against anchoring, not as paperwork
- Does the candidate raise false negatives unprompted, given nobody in the organisation is measuring them
- Whether candidate experience is argued as a source of signal quality and pipeline, not as politeness
Answer
Write the competency list before you name a single round
The question that produces a good loop is not "which interviews should we run" but "what will this person be doing eighteen months from now that nobody on the team can do today". Answer that concretely and the competencies fall out of it. A senior engineer joining a payments team mid-migration needs to design against a legacy system they did not build, land large changes incrementally without a freeze, debug production under time pressure, and raise the people around them by review and pairing rather than by writing all the code. Four competencies, and specific to this role — the same title on a greenfield internal-tools team yields a different list, weighted towards ambiguity tolerance and product judgement.
Most loops are built the other way round. Someone inherits four rounds from the last hire, each interviewer runs the question they enjoy, and the loop measures whatever those questions happen to measure. The failure is not that the questions are bad; it is that nobody can say which decision each round informs, so no round can be dropped, added, or fixed. Keep the list to four or five. A list of nine is not thoroughness, it is an admission that the role is undefined, and it guarantees each item is graded on a fragment of evidence.
One signal per round, proved by a coverage matrix
Draw a grid with competencies down the side and rounds across the top, and mark which round is the primary reader of which competency. Two things should be visibly true: every competency has an owner, and no round is the primary reader of something another round already owns.
| Round | Primary signal | Deliberately not scored here |
|---|---|---|
| System design, 60 min | Design against existing constraints | Coding fluency |
| Practical coding on real-ish code, 75 min | Incremental delivery, testing instinct | Algorithmic novelty |
| Debugging a seeded failure or a past incident walkthrough | Judgement under incomplete information | Design breadth |
| Collaboration and influence, 45 min | Raising others, handling disagreement | Technical depth |
The "deliberately not scored" column is the part teams skip and the part that does the work. Without it the coding round quietly grades design, the design round quietly grades communication, and you get four correlated opinions about one underlying impression. A loop where three rounds all reward fast pattern-matching on algorithms has not measured four things; it measured one thing four times and wasted three-quarters of the candidate's day and your team's. Duplicating a read is fine when chosen deliberately, for the competency hardest to see and most expensive to miss, but it belongs in the matrix rather than happening by accident.
Structured means the question is fixed and the probing is not
Structure is the largest single predictor of whether a loop tells you anything. In practice it means every candidate for the role gets the same core prompt in each round, in the same order, with the same opening framing. What varies is which follow-up thread the interviewer chases and how deep they go. Interviewer freedom moves from question selection to probing, which is where it belongs: that is where experience helps and where it cannot set a different bar for each candidate.
The rubric has to be anchored in observable behaviour rather than adjectives. A one-to-four scale with no anchors produces two interviewers using the same number to mean different things, which is worse than no score at all because it looks comparable.
Designs against existing constraints
1 — Proposes a greenfield design; never asks what exists today. 2 — Asks about the current system, then designs around it rather than through it; no migration path. 3 — Elicits the real constraints, names two or three trade-offs, sketches a path that ships in stages behind a flag. 4 — All of that, plus surfaces a constraint the interviewer had not mentioned and changes the design in response.
Anchors like these let a new interviewer be useful in week one, and they turn "the bar has slipped" from a feeling into a checkable claim.
Written feedback lands before anyone opens their mouth
Every interviewer submits written feedback — a score per competency, evidence quoted from the interview, a hire recommendation — before they see anyone else's and before the debrief starts, and the tool should lock it. This is not administrative hygiene; it is the only defence against the two effects that reliably destroy a debrief. The first is anchoring: whoever speaks first, or whoever is most senior, sets a frame the room adjusts towards rather than reasons from. The second is the cascade, where a genuine 3 becomes a 2 once the person holding it has heard two 2s, and the group ends up more confident than any individual was on less evidence than any individual had.
Written-first inverts the purpose of the debrief. It stops being a place where consensus is manufactured and becomes a place where disagreement is examined. When two interviewers land on 4 and 2 for the same competency, that divergence is the most informative thing in the room: usually they saw different behaviour, occasionally one of them is grading something the rubric does not describe, and now and then the rubric is wrong. All three are worth the ten minutes, and none of them survive a room that converged before the first person finished speaking.
Calibration is a separate, ongoing job: new interviewers shadow, then reverse-shadow with their scores compared against the lead's, and periodically the whole panel grades one recorded interview independently to see how far apart they have drifted.
The rejection you will never hear about
Regretted hires are visible. They sit near you, they surface in performance conversations, and each one makes the room a little more cautious next quarter. Strong candidates you rejected are invisible: hired by a competitor, promoted twice, never appearing in any metric you own. The organisation feels one error class and is structurally blind to the other, so loops drift in the only direction that feedback permits — tighter, more defensive, more willing to read absence of evidence as evidence of absence.
The asymmetry argument for caution is real; a bad senior hire costs a year of a team's tolerance and some of your credibility. But it argues for a well-designed loop rather than an unbounded one, and it stops working once you are rejecting people who would have succeeded. You can build partial visibility without ever knowing those outcomes: watch how often one interviewer's veto overturns an otherwise-positive loop, compare no-hire rates per interviewer against the panel, and audit how often "I got no signal" was recorded as a no-hire when the honest reading is that the round failed to create the opportunity.
Candidate experience is an input, not a courtesy
A senior engineer with options is running their own evaluation, and your loop is the highest-fidelity sample of your engineering culture they will see before signing. A round starting fifteen minutes late with an interviewer who has not read the CV tells them how the team treats preparation. A four-day unpaid take-home for a staff hire tells them what the team thinks other people's time is worth. Time from onsite to decision reads as a proxy for how the organisation makes any decision.
The self-interested reason matters more than the reputational one: a bad experience corrupts your data. A candidate who spent an hour defending themselves against a combative interviewer, or reverse-engineering a deliberately vague prompt, has produced no usable evidence about how they work with colleagues who are on their side. You cannot separate the low score from the conditions that caused it, so the round is unscoreable and the loop is one signal short. Interviewers prepared and on time, the shape of the loop explained in advance, a real block for the candidate's own questions with someone who will answer honestly, and a decision inside a stated number of days: these pay for themselves in signal quality before you count a single referral.
A loop is a measuring instrument, and two questions tell you whether yours works: can you name what each round measures, and can you name what it deliberately does not. Rubrics, written-first feedback and calibration all exist for one purpose — stopping four interviewers' impressions from collapsing into a single confident opinion that nobody can audit.
Likely follow-ups
- Two interviewers give the same candidate a 4 and a 2 on the same competency. What do you do in the debrief, and what do you do afterwards?
- How do you calibrate a brand-new interviewer without either wasting a real candidate or letting them grade one?
- Your loop has a 90% no-hire rate at the onsite stage. Is that a good bar or a broken loop, and how would you tell?
- How does this loop change if you are hiring the first senior engineer into a four-person team rather than the twentieth into a mature org?
Related questions
- You have three senior roles to fill this quarter. How do you build the pipeline and close the people you want?hardAlso on hiring6 min
- How do you take a mid-level engineer and get them to senior?mediumAlso on calibration6 min
- The domain expert tells you one thing and the written procedure says another. How do you work out which one the system should follow?hardSame kind of round: scenario5 min
- Your POC met every exit criterion and the customer bought from someone else. What went wrong?hardSame kind of round: scenario5 min
- You are convinced the company should walk away from a deal that everybody else wants to win. How do you make that case, and what do you do if you lose the argument?hardSame kind of round: scenario5 min
- In the room, the customer tells you a competitor has committed to something you cannot match. How do you respond?hardSame kind of round: scenario5 min
- A customer returns three failed units built in different weeks. You have one shift to say what else is affected. How do you bound it?hardSame kind of round: scenario6 min
- How would you price a feature that no customer asked for?hardSame kind of round: case-study6 min