Skip to content
QSWEQB
hardBehaviouralCase StudyScenarioLeadStaff

How would you design the interview loop for a senior engineer on your team?

Work backwards from the four or five competencies the role genuinely needs, map each to exactly one round so nothing is measured four times, score against shared behavioural anchors, require written feedback before the debrief opens, and treat both false negatives and candidate experience as data you owe yourself.

6 min readUpdated 2026-07-26

What the interviewer is scoring

  • Does the candidate derive the rounds from a competency list, or assemble a loop from the interviews their team already likes running
  • Whether the loop has a coverage matrix with no competency measured twice by accident and none left uncovered
  • That they can explain what a rubric anchor looks like, rather than defending a bare one-to-four scale
  • Whether independent written feedback before the debrief is presented as a mechanism against anchoring, not as paperwork
  • Does the candidate raise false negatives unprompted, given nobody in the organisation is measuring them
  • Whether candidate experience is argued as a source of signal quality and pipeline, not as politeness

Answer

Write the competency list before you name a single round

The question that produces a good loop is not "which interviews should we run" but "what will this person be doing eighteen months from now that nobody on the team can do today". Answer that concretely and the competencies fall out of it. A senior engineer joining a payments team mid-migration needs to design against a legacy system they did not build, land large changes incrementally without a freeze, debug production under time pressure, and raise the people around them by review and pairing rather than by writing all the code. Four competencies, and specific to this role — the same title on a greenfield internal-tools team yields a different list, weighted towards ambiguity tolerance and product judgement.

Most loops are built the other way round. Someone inherits four rounds from the last hire, each interviewer runs the question they enjoy, and the loop measures whatever those questions happen to measure. The failure is not that the questions are bad; it is that nobody can say which decision each round informs, so no round can be dropped, added, or fixed. Keep the list to four or five. A list of nine is not thoroughness, it is an admission that the role is undefined, and it guarantees each item is graded on a fragment of evidence.

One signal per round, proved by a coverage matrix

Draw a grid with competencies down the side and rounds across the top, and mark which round is the primary reader of which competency. Two things should be visibly true: every competency has an owner, and no round is the primary reader of something another round already owns.

RoundPrimary signalDeliberately not scored here
System design, 60 minDesign against existing constraintsCoding fluency
Practical coding on real-ish code, 75 minIncremental delivery, testing instinctAlgorithmic novelty
Debugging a seeded failure or a past incident walkthroughJudgement under incomplete informationDesign breadth
Collaboration and influence, 45 minRaising others, handling disagreementTechnical depth

The "deliberately not scored" column is the part teams skip and the part that does the work. Without it the coding round quietly grades design, the design round quietly grades communication, and you get four correlated opinions about one underlying impression. A loop where three rounds all reward fast pattern-matching on algorithms has not measured four things; it measured one thing four times and wasted three-quarters of the candidate's day and your team's. Duplicating a read is fine when chosen deliberately, for the competency hardest to see and most expensive to miss, but it belongs in the matrix rather than happening by accident.

Structured means the question is fixed and the probing is not

Structure is the largest single predictor of whether a loop tells you anything. In practice it means every candidate for the role gets the same core prompt in each round, in the same order, with the same opening framing. What varies is which follow-up thread the interviewer chases and how deep they go. Interviewer freedom moves from question selection to probing, which is where it belongs: that is where experience helps and where it cannot set a different bar for each candidate.

The rubric has to be anchored in observable behaviour rather than adjectives. A one-to-four scale with no anchors produces two interviewers using the same number to mean different things, which is worse than no score at all because it looks comparable.

Designs against existing constraints

1 — Proposes a greenfield design; never asks what exists today. 2 — Asks about the current system, then designs around it rather than through it; no migration path. 3 — Elicits the real constraints, names two or three trade-offs, sketches a path that ships in stages behind a flag. 4 — All of that, plus surfaces a constraint the interviewer had not mentioned and changes the design in response.

Anchors like these let a new interviewer be useful in week one, and they turn "the bar has slipped" from a feeling into a checkable claim.

Written feedback lands before anyone opens their mouth

Every interviewer submits written feedback — a score per competency, evidence quoted from the interview, a hire recommendation — before they see anyone else's and before the debrief starts, and the tool should lock it. This is not administrative hygiene; it is the only defence against the two effects that reliably destroy a debrief. The first is anchoring: whoever speaks first, or whoever is most senior, sets a frame the room adjusts towards rather than reasons from. The second is the cascade, where a genuine 3 becomes a 2 once the person holding it has heard two 2s, and the group ends up more confident than any individual was on less evidence than any individual had.

Written-first inverts the purpose of the debrief. It stops being a place where consensus is manufactured and becomes a place where disagreement is examined. When two interviewers land on 4 and 2 for the same competency, that divergence is the most informative thing in the room: usually they saw different behaviour, occasionally one of them is grading something the rubric does not describe, and now and then the rubric is wrong. All three are worth the ten minutes, and none of them survive a room that converged before the first person finished speaking.

Calibration is a separate, ongoing job: new interviewers shadow, then reverse-shadow with their scores compared against the lead's, and periodically the whole panel grades one recorded interview independently to see how far apart they have drifted.

The rejection you will never hear about

Regretted hires are visible. They sit near you, they surface in performance conversations, and each one makes the room a little more cautious next quarter. Strong candidates you rejected are invisible: hired by a competitor, promoted twice, never appearing in any metric you own. The organisation feels one error class and is structurally blind to the other, so loops drift in the only direction that feedback permits — tighter, more defensive, more willing to read absence of evidence as evidence of absence.

The asymmetry argument for caution is real; a bad senior hire costs a year of a team's tolerance and some of your credibility. But it argues for a well-designed loop rather than an unbounded one, and it stops working once you are rejecting people who would have succeeded. You can build partial visibility without ever knowing those outcomes: watch how often one interviewer's veto overturns an otherwise-positive loop, compare no-hire rates per interviewer against the panel, and audit how often "I got no signal" was recorded as a no-hire when the honest reading is that the round failed to create the opportunity.

Candidate experience is an input, not a courtesy

A senior engineer with options is running their own evaluation, and your loop is the highest-fidelity sample of your engineering culture they will see before signing. A round starting fifteen minutes late with an interviewer who has not read the CV tells them how the team treats preparation. A four-day unpaid take-home for a staff hire tells them what the team thinks other people's time is worth. Time from onsite to decision reads as a proxy for how the organisation makes any decision.

The self-interested reason matters more than the reputational one: a bad experience corrupts your data. A candidate who spent an hour defending themselves against a combative interviewer, or reverse-engineering a deliberately vague prompt, has produced no usable evidence about how they work with colleagues who are on their side. You cannot separate the low score from the conditions that caused it, so the round is unscoreable and the loop is one signal short. Interviewers prepared and on time, the shape of the loop explained in advance, a real block for the candidate's own questions with someone who will answer honestly, and a decision inside a stated number of days: these pay for themselves in signal quality before you count a single referral.

A loop is a measuring instrument, and two questions tell you whether yours works: can you name what each round measures, and can you name what it deliberately does not. Rubrics, written-first feedback and calibration all exist for one purpose — stopping four interviewers' impressions from collapsing into a single confident opinion that nobody can audit.

Likely follow-ups

  • Two interviewers give the same candidate a 4 and a 2 on the same competency. What do you do in the debrief, and what do you do afterwards?
  • How do you calibrate a brand-new interviewer without either wasting a real candidate or letting them grade one?
  • Your loop has a 90% no-hire rate at the onsite stage. Is that a good bar or a broken loop, and how would you tell?
  • How does this loop change if you are hiring the first senior engineer into a four-person team rather than the twentieth into a mature org?

Related questions

interview-loophiringstructured-interviewingcalibrationrubrics