Skip to content
QSWEQB

Agile and Scrum fundamentals

The answers a scrum master or delivery manager loop assumes: what each ceremony is for and how it fails, points as relative size, velocity as a forecast not a score, and dependencies reduced rather than managed. Fifty-nine items, sixteen worked through with a table, a forecast or a diagram.

59 questions

Go deeper on Agile, Scrum & Delivery

What agile was reacting to

What was the Agile Manifesto actually reacting against?

A specific industrial practice, not documentation or planning in general. In 2001 the dominant model was a sequenced project: a requirements phase producing a signed specification, a design phase producing a document set, then build, then an integration and test phase at the end, with change controlled through a formal board because change was assumed to be a defect in the original analysis. The observed failure was that the specification was wrong by the time it was built, and nobody found out for a year. Everything in the manifesto is a response to that single feedback-latency problem, which is why the useful reading of it is about shortening the loop between a decision and evidence about the decision. Read as a statement about disliking paperwork it becomes an excuse, and that misreading is where most of the subsequent damage came from.

Show me the four values with what each was reacting against.

The values are quoted constantly and their second halves are dropped, which is where the misuse begins.

value                       reacting against            how it is misused now
--------------------------  --------------------------  -------------------------
Individuals and inter-      staffing plans that         "we do not need process".
actions over processes      treated engineers as        Removing all structure and
and tools                   interchangeable resources   calling the result agile.
                            slotted into a defined
                            process

Working software over       a year of documents, a      "we do not write anything
comprehensive               signed spec, and no         down". Losing the design
documentation               running system until        record and the decisions,
                            month nine                  then relearning them

Customer collaboration      a contract that fixed       "the customer is always
over contract               scope, price and date up    available", when nobody
negotiation                 front, so both sides        has actually secured a
                            spent the project           decision-maker's time
                            defending their position

Responding to change        a change control board      "no plan". Reacting to
over following a plan       that treated new            whatever arrived last
                            information as a threat     week, with no thesis about
                            to the schedule             where the product is going

The sentence people forget:
  "That is, while there is value in the items on the right, we value the
   items on the left more."

The clause at the bottom is the whole of the answer to a hostile follow-up. The manifesto asserts a preference under conflict, not a prohibition, so a team that has stopped documenting anything has not applied the second value — it has deleted the right-hand side that the authors explicitly said had value.

The third row is the one worth dwelling on in an interview, because it is the value that fails for structural reasons rather than cultural ones. Collaboration over contract negotiation assumes a customer who can be present and can decide. In a fixed-price outsourced engagement, or with an internal stakeholder whose diary is full, that assumption is false, and every ceremony downstream of it — refinement, review, mid-sprint clarification — degrades quietly into the team guessing.

The fourth row is where the honest criticism lands. Responding to change without a plan is not agility, it is churn, and it is the most common lived experience of agile for engineers: a backlog reordered weekly, nothing finished, and no statement of what the product is for. The defensible position is that you keep a direction and hold the increments loosely, which is a harder discipline than either extreme.

What is the difference between agile and Scrum?

Agile is a set of values and principles with no mechanics; Scrum is one lightweight framework that implements them with specific accountabilities, events and artefacts. The relationship is that Scrum is an instance, not a synonym, and several other instances exist — Kanban, Extreme Programming, and various in-house hybrids. This matters in an interview because the two questions have different answers: "are you agile" is about feedback loops and who decides, while "do you do Scrum" is about whether you have a product owner, a sprint and a retrospective. A team can run every Scrum event faithfully and be entirely un-agile if the backlog is dictated, the increment is never shown to a user, and the retrospective produces nothing. The reverse also happens, and a team shipping daily against real user feedback with no sprints at all is the stronger position to defend.

Why do the twelve principles do more work than the four values?

Because the principles are specific enough to be falsified, and the values are not. "Deliver working software frequently, from a couple of weeks to a couple of months, with a preference to the shorter timescale" makes a claim you can check against a team's release history; "individuals and interactions over processes" does not. The principles are also where the parts everyone skips live: a sustainable pace indefinitely, technical excellence and good design enhancing agility, the best architectures emerging from self-organising teams, and simplicity defined as maximising the work not done. Those four are the ones a team under pressure abandons first, and they are the ones that make the rest possible. Quoting a principle rather than a value is also a quick signal in an interview that you have read past the poster.

What separates a team that is agile from one that performs agile?

Where decisions are made and how fast evidence arrives. In a genuinely agile team the people doing the work decide how it is done, the increment reaches someone who will use it inside a few weeks, and what is learned changes the next increment. In a performing team the events happen on the calendar, the board is tidy, the artefacts exist, and the sequence of decisions is unchanged from the plan-driven model — scope is handed down, the date is fixed, and the sprint is a reporting interval. The cheapest diagnostic is to ask what the team changed last month because of something it learned. A team that cannot answer has ceremonies without feedback, which is the expensive combination: all of the meeting cost and none of the adaptation benefit.

Why does agile have a bad reputation among engineers?

Because most engineers have met it as an increase in surveillance rather than autonomy. The lived version is a daily meeting where you justify yesterday to a manager, estimates that become commitments and then become deadlines, points tracked as productivity, a backlog reordered without explanation, and a retrospective that produces the same three items every fortnight. Every one of those inverts what the practice was for, and the inversion is usually not malicious — it happens because the artefacts of agile are legible to management while the substance is not, so the visible parts get adopted and measured. The honest thing to say in an interview is that the cynicism is earned, and that defending the framework rather than acknowledging it is the fastest way to lose an engineering audience.

When is agile the wrong answer?

When the cost of an increment is very high or the requirement genuinely cannot change. Certification-bound work where each release triggers a regulatory submission, hardware with a fixed tape-out, or a contractual integration against a frozen external specification all have feedback loops measured in quarters, and iterating faster than the loop returns evidence just produces churn. The nuance worth adding is that this is a statement about release cadence rather than about working practice: an avionics team can still work in short internal iterations, keep the build integrated continuously and hold retrospectives, while releasing annually. So the wrong answer is dogmatic Scrum with a two-week sprint review nobody can attend, and the right one is picking the loop length from the domain rather than from the framework.

Scrum mechanics

What are the three Scrum accountabilities, and what is each actually responsible for?

A product owner, a scrum master and developers, and the deliberate word is accountability rather than role. The product owner is accountable for the value of the product and, concretely, for the ordering of the backlog — one person deciding, so that "everything is priority one" cannot survive. The developers are accountable for the increment: how the work is done, the technical decisions, the sprint plan, and honouring the definition of done. The scrum master is accountable for the effectiveness of the team and of Scrum in the organisation, which is coaching, facilitation and removing impediments including those above the team. The line that gets tested is that none of these three is a manager of the others, and that the developers own the how absolutely — a product owner who assigns tasks or dictates the sprint plan has taken an accountability that is not theirs.

Why is a scrum master not a project manager?

Because they have no authority over scope, date, budget or people, and no accountability for delivery in the sense a project manager carries it. A project manager plans the work, assigns it, tracks it against a baseline and reports variance; a scrum master makes the team capable of doing that for itself and works on the system that keeps getting in the way. The practical distinction to name is direction of effort: a project manager's day is spent on the plan, a scrum master's on the impediments, and if the scrum master is maintaining the plan then the team has a project manager with a different title. The honest caveat is that many advertised scrum master jobs are delivery-management jobs, and the useful interview move is to ask which one this is rather than pretend the distinction is always respected.

What does a product owner own that nobody else can?

The order of the backlog, and the accompanying right to say no. Ordering is a single-person accountability by design, because value trade-offs made by a committee produce a list with no order at all, and a team facing an unordered list either picks for itself or takes whatever was shouted loudest. Beyond ordering, the product owner owns the product goal, the reasoning behind the priority so the team can make sensible local decisions, and availability during the sprint to answer the questions refinement did not anticipate. The failure mode is the proxy product owner: someone with the title, no mandate, and a stakeholder group that overrides them, which converts every prioritisation conversation into an escalation and is the root cause behind most stories that arrive half-specified.

Show me the five Scrum events with the timebox and the failure mode of each.

Naming the events is a screening question; naming how each one rots is the answer that distinguishes someone who has run them.

event            timebox        purpose                  how it fails
---------------  -------------  -----------------------  ----------------------
Sprint           <= 1 month     a container with a goal  becomes a reporting
                 fixed length   that yields a usable     interval. Length varies
                                increment                to fit the work, so no
                                                         forecast is possible

Sprint planning  8 hrs for a    why this sprint matters, planning becomes task
                 1-month        what is taken, how it    assignment by the PO or
                 sprint         will be done             the lead. No sprint goal
                                                         is written at all

Daily scrum      15 min, for    Developers re-plan the   status report to the
                 the             day toward the sprint    manager. Round the room,
                 Developers      goal                     three questions, nobody
                                                          listens to anyone else,
                                                          and it runs to 30 mins

Sprint review    4 hrs for a    inspect the increment    a demo with no
                 1-month        with stakeholders and    stakeholders present. It
                 sprint         adapt the backlog        becomes sign-off theatre
                                                         and the backlog does not
                                                         change as a result

Sprint retro-    3 hrs for a    inspect how the team     no actions, or actions
spective         1-month        worked and commit to     with no owner that recur
                 sprint         one or two improvements  every fortnight. Or it is
                                                         cancelled first when the
                                                         sprint is late

Refinement is not an event in the Guide. It is ongoing work, which is why
"we do not have time for refinement this sprint" is a real answer and
"we skipped the retro" is not.

Say "the Developers" rather than "the dev team" here, because the vocabulary dates you. The 2020 Scrum Guide removed the development team as a separate unit inside the Scrum Team and made Developers an accountability instead, and it is explicit that if the Product Owner or Scrum Master is actively working on items in the Sprint Backlog, they participate in the Daily Scrum as Developers. So the event is not closed to those two roles by title; it is for whoever is doing the work the sprint goal depends on, and anyone else present is an observer.

The daily scrum row carries the most weight in an interview because its failure is so nearly universal. The test is who the meeting is addressed to: if each person speaks to the scrum master or the manager, it is a status report and the information flows out of the team rather than around it. The fix is not banning managers, it is changing the unit of the conversation from people to work — walk the board right to left, discuss items rather than individuals, and ask what is between us and the sprint goal.

The review's failure mode is the most expensive, because the event exists to generate the feedback the whole framework depends on. A review with no stakeholders in the room produces no adaptation of the backlog, which means the team is iterating on delivery and not on the product. If stakeholders will not come, that is the impediment, and the scrum master's job is to fix attendance or to say plainly that the team is building without evidence.

The retrospective row is where the culture is visible. A retrospective cancelled because the sprint is behind states the organisation's actual priority: the mechanism for getting faster is the first thing sacrificed to going faster this fortnight.

Note that four of the five have a timebox that scales with sprint length and the daily does not. Fifteen minutes is absolute, and a daily scrum that reliably runs to thirty is not a discipline problem — it means detailed problem-solving has nowhere else to go, and the fix is a slot immediately afterwards for the two people who actually need it.

What is a sprint goal for, and what happens without one?

It is the single objective that makes the sprint's items a coherent bet rather than a batch, and its real function is to let the team make trade-offs without asking. With a goal, a developer discovering on Wednesday that one story is larger than thought can drop or simplify something and still deliver the outcome; without one, every item is equally mandatory and the only lever left is working longer. The absence has a recognisable signature: a sprint described as a list, a review that walks tickets rather than showing an outcome, and a partially-complete sprint that delivered nothing usable because the eight half-done items were unrelated. A goal also gives the product owner something to protect during the sprint, which is what makes refusing mid-sprint injection an argument about the goal rather than about the team's willingness.

Show me how work arriving mid-sprint is handled.

The wrong question is whether to accept it; the right one is what it displaces and who decides.

flowchart TD
    A[New request arrives mid-sprint] --> B{Does it threaten the sprint goal}
    B -->|no| C[Goes to the product backlog<br/>ordered by the product owner]
    B -->|yes| D{Is it worth more than what it displaces}
    D -->|no| C
    D -->|yes| E[Product owner names what comes out<br/>similar size, agreed with the developers]
    E --> F[Sprint goal restated<br/>or the sprint is cancelled]

The first decision node is the one teams skip. Most mid-sprint requests are not urgent, they are merely recent, and the default path for a recent request is the backlog — where the product owner will order it against everything else next week. Saying that out loud converts an interruption into an ordinary prioritisation decision, and it costs the requester nothing they were entitled to.

The exchange at the fourth node is the mechanism that makes the answer credible. Nothing is added without something of comparable size leaving, and the person who names what leaves is the product owner rather than the team, because it is a value decision. A team that absorbs additions without removals is not being helpful, it is quietly converting the sprint into overtime and then being blamed for the spillover.

Cancelling the sprint is a real option and candidates almost never mention it. The Guide gives that power to the product owner and it is correct precisely when the goal has become obsolete — a regulatory change, a competitor launch, a production incident that invalidates the increment. Continuing to work toward a goal nobody wants because the sprint has four days left is the worse outcome, and being willing to say so is a senior signal.

The genuinely hard version of this question is a team where injection happens weekly. That is not a facilitation problem, it is a capacity structure problem, and the durable answers are an explicit unplanned-work allowance sized from history, or a rota where one developer handles interrupts while the rest protect the goal. If more than roughly a third of capacity is unplanned, the honest recommendation is to stop sprinting and run the work as a flow system with WIP limits, because a sprint commitment against that variance is a fiction everyone is maintaining.

Show me a retrospective format, including the step most teams skip.

Formats vary and the omission does not, so the shape is worth having written.

60-75 minutes, fortnightly, developers plus PO and SM. No managers unless
the team invites them.

step                     min  what happens
-----------------------  ---  -----------------------------------------------
0  Check last time's      10  Read out the actions from the previous retro.
   actions                    State done / not done / abandoned, by name.
                              THIS IS THE STEP TEAMS SKIP.

1  Set the stage           5  Restate the prime directive. Confirm what is
                              in scope, and that this is not a review.

2  Gather data            15  Facts before feelings. Timeline of the sprint,
                              the metrics - cycle time, spillover, unplanned
                              work, incidents - written up beforehand.

3  Generate insight       20  Why did that happen. Group the observations,
                              then pick ONE theme by dot vote rather than
                              discussing all seven.

4  Decide what to do      15  One or two actions maximum. Each with a named
                              owner, a size small enough to fit in the next
                              sprint, and a place on the board so it is
                              visible work rather than goodwill.

5  Close                   5  How was this retro. One word each.

Rotate the facilitator. A scrum master who facilitates every retro for a
year owns the team's improvement instead of the team owning it.

Step zero is the entire difference between a retrospective that changes something and a fortnightly complaint session. Without it, actions are generated enthusiastically and never referenced again, so the same three items recur for six months and the team correctly concludes the meeting is theatre. Reading the previous actions out loud, with owners named, makes non-delivery visible and therefore uncomfortable, which is the only mechanism that makes the next set of actions real.

The cap of one or two actions is the second thing teams get wrong. A retro producing nine improvements delivers none of them, because they compete with the sprint for the same capacity and lose. One action that lands beats nine that are logged, and the arithmetic is worth stating: one improvement per fortnight is twenty-six a year, which is far more change than any team actually absorbs.

Putting the action on the board is the detail that converts intent into throughput. Improvement work that lives only in the retro notes is being funded from slack that does not exist. If the team cannot put it on the board, the honest conclusion is that there is no capacity for improvement, and that is the finding to escalate rather than to absorb.

The facts-before-feelings ordering matters because a retro opened with "how did that feel" anchors on whoever speaks first and loudest. Bringing the sprint's numbers already prepared — where items sat, what was blocked and for how long — gives the quieter half of the team something to point at.

Show me a sprint that will not complete, handled as options.

Halfway through, it is clear the sprint goal is at risk. The answer is not a status colour, it is a set of priced choices, raised on Wednesday rather than on the last day.

Sprint 41. Goal: "a customer can check out with a saved card."
Day 6 of 10. Remaining: 5 of 9 stories, including the two largest.
Cause: the payment provider's sandbox was unavailable for three days,
and the tokenisation story turned out to need a schema change.

OPTION A  Cut scope, keep the goal
  do        drop the saved-card management screen and the receipt email.
            Checkout with a saved card still works end to end.
  cost      customers cannot delete a saved card until next sprint.
            Support will field a handful of tickets.
  gains     the goal is met, the increment is releasable, the forecast
            for the next sprint is undamaged.
  who       product owner decides. My recommendation.

OPTION B  Keep scope, extend the sprint
  cost      breaks the fixed cadence, so velocity and cycle time become
            incomparable with previous sprints and every forecast
            degrades. Next sprint's planning slips too.
  gains     nothing that A does not give, one sprint later.
  who       nobody. This is the option to name and refuse.

OPTION C  Keep scope, work longer hours
  cost      about 15 extra hours across four people. Historically this
            team's defect rate roughly doubles in a crunch sprint, so
            the work returns next sprint as rework.
  gains     the appearance of the commitment being met.
  who       refuse. If asked to authorise it, say what it costs.

OPTION D  Carry the two large stories, deliver nothing
  cost      no releasable increment. The stakeholder learns of the slip
            at the review, four days from now.
  gains     none. This is the default outcome if no decision is made,
            which is why the decision is needed today.

What I need: a decision between A and the consequence of D, from the PO,
today. And the sandbox dependency goes on the impediment log as the
actual cause, because it will recur.

The load-bearing feature is the timing. Raising this on day six leaves options A and D genuinely open; raising it on day ten leaves only D, dressed as a surprise. Interviewers listen for whether you report a forecast change when the probability moves rather than when the outcome is certain, because that is the difference between managing delivery and narrating it.

Option B is included because it will be suggested and it is the one to refuse on technical grounds rather than dogmatic ones. A variable-length sprint destroys the only property that made the cadence useful for forecasting — comparability — so the cost is not "the Guide says no", it is that every subsequent estimate loses its evidence base.

Option C deserves the explicit number. Naming a doubled defect rate from the team's own history converts an argument about commitment and attitude into an arithmetic one about where the work reappears, and it is much harder to override.

The last paragraph is what most candidates omit: separating the recovery from the cause. Cutting scope handles this sprint, and the provider sandbox will do it again next quarter unless somebody owns the dependency. A scrum master who fixes only the symptom has facilitated, not removed an impediment.

What is spillover really telling you?

Usually that the stories were too large or too coupled, not that the team was slow. A story spilling once is noise; a pattern of the same one or two large items carrying across sprints is a splitting problem, and the diagnostic is whether the spilled item was ever demonstrably at ninety per cent or simply opaque until it was done. The second common cause is invisible queueing — items finished by the developer and sitting in review or waiting on a test environment, which shows up as spillover and is actually a flow problem in a column nobody limits. What the metric should never be used for is a completion-rate score reported upward, because the reliable response to that pressure is smaller commitments and status inflation rather than better slicing.

The backlog

What is a product backlog, and what makes it different from a list of tasks?

It is a single ordered list of everything that might change the product, with exactly one owner of the order, and the two words doing the work are "single" and "ordered". Single, because the moment there are three lists — the roadmap, the bug tracker, the technical debt spreadsheet — nothing is genuinely prioritised against anything else and the team's real priority is whoever asked most recently. Ordered rather than categorised, because priority buckets always collapse into a large "high" bucket. The other distinction from a task list is granularity by distance: the top is refined, small and ready, and the bottom is deliberately coarse, because refining items that will never be built is inventory. A backlog where every item is fully specified is a specification with a different name.

What is a definition of ready, and when does it become harmful?

It is the team's agreement about what a story needs before it can be pulled into a sprint — typically a clear outcome, testable acceptance criteria, dependencies identified, a size the team believes fits, and no open question that would stall the work. Used as a guide it prevents the commonest waste, which is starting an item and discovering on day three that nobody can answer the central question. Used as a gate it becomes harmful, and the failure is recognisable: refinement turns into a formal handover, the product owner is asked to produce a document before anyone will discuss the problem, and the team acquires a legitimate excuse for not starting anything. The healthy version treats readiness as a conversation's outcome rather than a checklist's, and accepts that some uncertainty is best resolved by building a thin slice.

Show me a definition of done written concretely.

A definition of done is a shared, binary, testable statement of what "done" means for every item, and most teams have a vague one.

DEFINITION OF DONE - Checkout team, agreed 12 May 2026, reviewed quarterly

Applies to every product backlog item. Not negotiable per story.

Code
  [ ] merged to main behind a feature flag, no long-lived branch
  [ ] reviewed by someone who did not write it
  [ ] no new lint or type errors, no TODO referencing this story

Tests
  [ ] unit tests for new logic, and one covering the bug if this was a bug
  [ ] one end-to-end test for the primary happy path
  [ ] full suite green on main, not just on the branch

Operability
  [ ] logs and a metric emitted for the new path
  [ ] alert exists if the path can fail silently
  [ ] runbook entry updated if on-call behaviour changes

Deployed and observable
  [ ] running in production, flag off is acceptable
  [ ] verified in production by the developer, not by QA in staging

Documentation
  [ ] API reference regenerated if the contract changed
  [ ] a line in the release note that a support agent can understand

NOT in the DoD, deliberately
  - performance testing. Done per-epic, not per-story. Named here so its
    absence is a decision rather than an oversight.
  - accessibility audit. Per-release, with a specialist.

Weaker than we want, and known: security review is per-epic because we
do not have the reviewer capacity for per-story. Reviewed in October.

The property that makes this usable is that every line is binary and checkable by someone other than the author. "Code is well tested" is not a definition of done, it is an aspiration, and it produces the argument at review that the definition existed to prevent. If two people can disagree about whether a line is met, the line needs rewriting.

The deployment line is where most teams' definitions are quietly false. A definition of done ending at "merged" means the increment is not usable at the review, so the sprint produced no evidence — which is the whole purpose of the sprint. Including production, even behind an off flag, is what makes the review about the product rather than about a branch.

The two absences are the part an interviewer will notice. A definition of done that omits performance and accessibility silently is a team accumulating an unmeasured obligation; one that names them as handled elsewhere has made a trade-off you can interrogate. The same applies to the last paragraph: writing down the part that is weaker than you want, with a review date, is how a definition stays honest rather than becoming a document nobody believes.

The rule that makes it work at all is that it is not negotiated per story. The moment "done except the tests" is acceptable under deadline pressure, the definition has become advisory, and every later estimate is measuring a different quantity from the one before it.

Show me a horizontally sliced story rewritten as vertical slices.

Horizontal slicing is the default instinct because it follows the architecture, and it is the single most common reason a sprint delivers nothing usable.

HORIZONTAL - sliced by layer. Nothing works until all four land.

  1  "Create the saved_cards table and migration"          3 pts
  2  "Add the tokenisation service and provider client"     5 pts
  3  "Build the saved-card API endpoints"                   5 pts
  4  "Build the saved-card UI"                              3 pts

  what a stakeholder can see after story 3 of 4:  nothing
  what is releasable after story 3 of 4:          nothing
  if the sprint ends after story 3:               16 points of
                                                  inventory, zero value
  who can work in parallel:                       nobody, it is a chain

VERTICAL - sliced by outcome. Each one is a usable thin path.

  1  "A customer can pay with a card and tick 'save this card'.
      One card only. Shown as the last four digits at next checkout.
      Cannot be deleted yet."                              5 pts
      -> touches table, service, API and UI. Releasable. Demoable.

  2  "A customer can save up to five cards and choose which one
      to pay with."                                        3 pts

  3  "A customer can delete a saved card."                 2 pts

  4  "A customer whose saved card is declined is offered a
      one-time card entry without losing the basket."       3 pts

  after slice 1:   a real customer can do a real thing. Feedback exists.
  after slice 1:   the provider integration is proven, so the largest
                   unknown in the epic is resolved in week one.
  if the sprint ends after slice 3:  goal met, slice 4 carries cleanly.

The decisive difference is where the risk sits. The horizontal version defers the integration with the payment provider until story two of four and defers all evidence to the end, so the biggest unknown is discovered last. The vertical version drives one narrow path through every layer immediately, which is why the first slice is deliberately the least capable and the most valuable: it proves the architecture and produces something a stakeholder can react to.

The second difference is optionality. Vertical slices can be stopped at any point with a coherent product; horizontal slices have exactly one acceptable stopping point, which is the end. That is what makes a horizontally sliced sprint all-or-nothing and why its spillover is total rather than partial.

The objection to rehearse is that vertical slicing means touching the same code four times and is therefore wasteful. Sometimes true, and it is the right price: the rework is bounded and known, while the cost of building four layers against an unvalidated assumption is unbounded. If the rework genuinely dominates, the answer is a thin technical spike first, timeboxed, not a return to layer slicing.

The tell that a story is still horizontal is the language. A title beginning with "create", "add" or "build" and naming a component is a task; one describing an actor doing something and getting a result is a slice. "Set up the message queue" cannot be demonstrated to anyone, which is the test.

What are the story-splitting techniques worth knowing by name?

Six carry most of the work. Split by workflow step, taking the thinnest end-to-end path first and adding steps later. Split by business rule, implementing the common case before the exceptions and discounts. Split by variation in data or interface, one payment provider or one locale before the rest. Split by operation — view, then create, then edit, then delete — since read paths are usually independently valuable. Split by effort, deliberately isolating the uncertain part so it can be spiked. And split by quality attribute, shipping a correct but slow version before the optimised one, which is legitimate only if slow is genuinely acceptable to a user. The unifying test is that every resulting slice remains demonstrable to someone who does not know the architecture; if a slice can only be explained in terms of components, it is a task.

Where does technical debt live in a backlog?

In the same ordered list as everything else, described in terms of the cost it is imposing, because a separate list is a list that never gets scheduled. The framing that survives contact with a product owner is consequence rather than virtue: "deploys to this service take forty minutes and half of them need a manual step, which is where two of our last three incidents came from" is tradeable, while "we need to refactor the payments module" is not. Some debt work never wins that competition honestly and is better handled inside feature work as the definition of done, or by a standing allowance the team holds without asking. The position to avoid in an interview is asking for a dedicated debt sprint, which reliably gets cut and signals that the debt was not attached to any outcome.

What does refinement look like when it works?

A short, frequent conversation that ends with the team understanding the problem, not the product owner presenting solutions. Practically that is thirty to sixty minutes once or twice a week, on the items likely to be taken in the next sprint or two, with the developers asking the questions that reveal hidden scope — what happens when this fails, which customers have this state today, what does the existing data actually look like. The output is a smaller, clearer, sized item, and frequently the discovery that one story is three. The failure mode is refinement as estimation theatre: forty items walked through so numbers can be attached, no splitting, and no question that changes anything. The diagnostic to offer is how often refinement changes the shape of a story, because a session that never does is a meeting for recording estimates.

Estimation and forecasting

What is a story point actually measuring?

Relative size, which bundles complexity, volume of work and uncertainty into one comparative number — and deliberately not duration. The reason for relative rather than absolute is that humans are consistently poor at estimating time and considerably better at judging that one thing is about twice another, so a sequence of comparisons is more stable than a sequence of guesses. The consequence is that a point has no meaning outside one team, since it is calibrated against that team's own reference stories. Points earn their keep only when combined with measured throughput to produce a forecast; if a team estimates in points and never uses velocity to forecast anything, the estimation is ceremony and the honest recommendation is to count stories instead, which is cheaper and roughly as predictive.

Show me story point estimates as relative sizes, and the anchoring mistake.

Points only work as comparisons to something the team has already built, so the reference set is the whole technique.

REFERENCE STORIES - agreed by this team, revisited quarterly

  1 pt  "add a new field to the customer export"
        known pattern, one file, no unknowns

  2 pt  "add a validation rule to the address form,
         with the error state"

  3 pt  "a customer can delete a saved card"
        our 3 is the anchor. If unsure, ask: bigger or smaller than this

  5 pt  "a customer can save a card at checkout, one card,
         no management screen"
        touches four layers, one external call, some unknown

  8 pt  "migrate the address table to the new schema
         with a backfill"
        large, or genuinely uncertain. 8 is a warning, not a size

  13    do not estimate it. Split it or spike it.

THE ANCHORING MISTAKE

  Wrong:  the tech lead says "this is a 5, right?" and the team agrees.
          Six people now hold one person's estimate. The spread that
          contained the information has been destroyed.

  Wrong:  "5 points is about 2 days for us." Now every estimate is a
          duration in disguise, and optimism returns in full.

  Right:  everyone reveals simultaneously.
          Round 1:   2  3  3  5  13  3
          The 13 is the most valuable number in the room. Ask that
          person first, then the 2.
          The 13 knew the export runs through the legacy tax service.
          Round 2:   5  5  8  8  8  5   -> take 8, or split it.

The single-round-then-discuss ordering is the mechanism, not the etiquette. Simultaneous reveal exists to protect the outliers, and the outliers are the point: a wide spread means the team does not share an understanding of the story, which is more useful than any number. Estimating serially, or after the loudest person has spoken, converts six independent judgements into one, and that one is usually the most senior and least recently hands-on.

Asking the high estimate to speak first is a deliberate inversion. The person who sees more work usually knows something — a legacy dependency, a data problem, a migration — and surfacing it early is the actual product of the meeting. Asking the low estimate first tends to establish a floor that the high estimate then has to argue against.

The reference set is what stops the scale drifting. Without one, this quarter's 5 is last quarter's 3, velocity becomes meaningless as a forecasting input, and nobody notices because the number still goes up. Revisiting the anchors periodically, especially after the team changes shape, is the maintenance nobody mentions.

The 13 rule is worth stating as policy. Above a certain size an estimate carries no information — the team is really saying "we do not know" — so the correct output is a split or a timeboxed spike rather than a larger number. A backlog containing 20s and 40s is a backlog of unexamined work with decoration.

Why does converting points to hours defeat the purpose?

Because it reintroduces exactly the estimation error that relative sizing was adopted to avoid, and then adds false precision on top. Once a point is eight hours, a five-point story is a forty-hour promise, and the number stops being a comparison and becomes a schedule that can be checked, disputed and enforced. Two concrete harms follow. Uncertainty disappears: a five that meant "large and partly unknown" becomes a duration with no error bar, so the unknown is silently priced at zero. And the team is now measurable in hours, which invites capacity arithmetic — six people times eighty hours divided by eight — that plans to a hundred per cent of a quantity that was never a time in the first place. If management needs hours, give them a forecast range from throughput instead, which is a genuinely better answer to their actual question.

What is velocity for, and what is it not for?

It is an input to a forecast for one team, and it is a measure of nothing else. Used properly it answers "how many sprints is this remaining scope likely to take" by combining the team's recent range with the remaining estimate, and it is most honest expressed as a band from the worst few sprints to the best. It is not a productivity measure, because the quantity being counted is the team's own calibration and can be inflated at will by estimating the same work higher — a change nobody outside the team can detect. It is not a target either: velocity set as a goal rises reliably and delivery does not, which is Goodhart operating in the space of a single fortnight. The line to say aloud is that velocity is a capacity observation, and the moment it is reported as performance it stops being an observation.

Show me velocity used as a forecast range rather than a commitment.

The same eight sprints of history, read the way it defeats a candidate and then the way it survives a stakeholder.

Team: Checkout. Completed points per sprint, last 8 sprints:
    32, 24, 41, 19, 36, 28, 22, 38
    mean 30   median 30   worst 19   best 41

Remaining scope for the saved-cards epic: 118 points estimated,
plus 6 items not yet estimated - call it 130 to 145 in practice.

THE ANSWER THAT BECOMES A SLIP
    118 / 30 = 3.9 sprints  ->  "4 sprints, so 22 September"
    Uses the mean, ignores the unestimated items, ignores that two
    of eight sprints came in at or below 22 and one of those was 19.
    Roughly a coin flip at best.

THE ANSWER TO GIVE
    pessimistic  130 / 19 = 6.8  ->  7 sprints  ->  3 November
    likely       130 / 30 = 4.3  ->  5 sprints  ->  6 October
    optimistic   130 / 41 = 3.2  ->  4 sprints  ->  22 September

    "Four to seven sprints, most likely five. I would plan the launch
     against 3 November and I would not announce 22 September."

WHY THE RANGE IS THAT WIDE
    - the 19 sprint was two people on an incident for a week. That
      recurs about once a quarter, so it belongs in the range.
    - 6 items are unestimated. Scope has historically grown ~15%
      between refinement and done on this epic.
    - one dependency on the platform team, not yet scheduled. This is
      the single item that could move the far end past November.

WHAT NOT TO DO
    - do not average away the bad sprints. They are the variance, and
      the variance is the forecast.
    - do not re-estimate to make the number fit the date. That changes
      the units and destroys the history you are forecasting from.

The essential move is to forecast from the observed spread rather than from the mean, because the mean encodes an assumption that nothing will go wrong — and the history in front of you says something goes wrong roughly every fourth sprint. The bad sprints are not outliers to be excluded; they are the sample of how this team actually behaves in this organisation.

Naming the width's causes is what converts a range into a conversation. Two of the three reasons here are actionable by someone other than the team: the platform dependency can be scheduled and the unestimated items can be refined. Presenting the range without them reads as hedging, and presenting them turns the stakeholder into a participant in narrowing it.

The last block is the failure interviewers probe. Re-estimating to fit a date is common, feels like negotiation, and silently recalibrates the unit — after which velocity is no longer comparable to its own history and the team has lost the only evidence base it had. If pressure requires a smaller number, the honest lever is scope, not the scale.

The sentence worth rehearsing is the recommendation with a stated risk appetite: plan against the pessimistic end for anything externally announced, against the likely end for internal sequencing. That hands the risk decision to the person who owns it rather than absorbing it silently and apologising in October.

Why is comparing velocity between two teams meaningless?

Because a point is defined by each team's own reference stories, so the two numbers are denominated in different units and the comparison is arithmetic performed on incompatible scales. Team A calling a certain piece of work five points and Team B calling it two says nothing about either team's output; it says their anchors differ. The harm is not merely that the comparison is invalid but that publishing it changes behaviour in exactly the wrong direction: the team being compared unfavourably inflates its estimates, which is undetectable from outside, costless, and immediately effective. If a genuine cross-team comparison is needed, the only defensible measures are unit-free — throughput of items, cycle time, change failure rate, deployment frequency — and even those need context on the work's nature before they mean anything.

Show me a Monte Carlo forecast as a probability table of dates.

Throughput sampling replaces the estimate entirely, which is why it is the strongest forecasting answer available in an interview.

Input: completed ITEMS per sprint, last 12 sprints. No points involved.
    9, 7, 12, 5, 11, 8, 10, 6, 13, 9, 4, 11

Remaining: 47 items. Scope growth allowance: historically about 15 new
items appear for every 100 completed on this epic, a multiplier of
1.15, so simulate 47 x 1.15 = 54.

Method: 10,000 trials. Each trial draws sprints at random with
replacement from the 12 observed values until 54 items are done, and
records how many sprints that took. Sort the 10,000 results.

The output is a chance per date, not a menu of confidence levels. Read
down the column, do not pick a percentage and look up a date.

  sprints   date      chance of being done by   what it is for
  --------  --------  -----------------------   ---------------------
     5      29 Sep              5%              nothing. Wishful.
     6      13 Oct             45%              worse than a coin flip
     7      27 Oct             86%              internal sequencing
     8      10 Nov             98%              what I commit to
                                                externally
     9      24 Nov            >99%              a contractual or
                                                regulatory date, or a
                                                marketing launch

  Read as: "10 November is a 98% date. 13 October is a 45% date, which
  is another way of saying it will probably be late."

WHY THIS BEATS points / velocity
  - no estimation step, so no estimation error and no anchoring
  - the 4-item and 5-item sprints are in the sample, so holidays,
    incidents and review queues are already priced in
  - it outputs a distribution, so the risk decision belongs to the
    person who owns the risk

WHAT IT DOES NOT FIX
  - item size drift. If the team starts slicing smaller, throughput
    rises with no change in delivery. Sanity-check against dates.
  - a scope change that adds 30 items. No model absorbs that.
  - fewer than ~8 sprints of history is too small a sample to
    resample from. Say so rather than producing a table anyway.

The reason this is more honest than a velocity division is that the input is measured rather than estimated. Twelve sprints of throughput already contain every tax on the team — leave, on-call, a slow review queue, the fortnight two people spent on an incident — whereas an estimate of remaining effort implicitly assumes those away and then discovers them one at a time.

Counting items rather than points is the part candidates find surprising. Item counts forecast about as well as point sums for most teams, because story sizes cluster, and they remove the whole estimation apparatus along with its gameability. If a team is sceptical, run both for three sprints and compare — the forecasts usually agree, and one of them is free.

The scope growth multiplier is where most forecasts quietly fail. Backlogs grow as work is understood, and a model of the currently-known items answers a question nobody asked. Measuring the historical growth rate and applying it makes the assumption visible and arguable, which is better than a footnote saying scope is assumed stable.

The limitations block is what makes the answer credible rather than a sales pitch for a technique. Volunteering that item-size drift can inflate throughput, and that a short history is not a sample, demonstrates that you understand the model rather than having read about the tool.

How do you plan capacity without planning to a hundred per cent?

By starting from measured historical throughput rather than from available hours, and by leaving explicit room for the work that always arrives. Planning to full capacity fails for a structural reason rather than a motivational one: it assumes zero variance, zero unplanned work and no queueing, and a system loaded to a hundred per cent has no slack to absorb any of the three, so wait times grow sharply and everything is late together. Practically, take the recent range rather than the best sprint, subtract known absence, hold back a named allowance for production support and interrupts sized from the last few sprints rather than from optimism, and count the improvement action from the retrospective as real work. The argument to make to a stakeholder is that slack is what makes dates predictable, so removing it buys apparent commitment at the price of the forecast.

What is planning poker actually for?

Surfacing disagreement, not producing a number. The number is a by-product; the value is in the second where six people reveal different estimates and the team discovers that half of them did not know about the legacy dependency, or that two people are estimating different scopes because the acceptance criteria are ambiguous. That is why the mechanics are simultaneous reveal, outliers speak first, and a second round after the discussion — every element exists to protect the spread from being collapsed by the most senior voice. It follows that a team whose estimates always agree on the first round is either genuinely calibrated or, far more often, deferring to someone, and the fix is to check whether anyone can articulate the risks. When the discussion has happened, the estimate itself is almost incidental.

Flow and Kanban

What does Kanban change relative to Scrum?

It removes the timebox and the batch, and replaces the sprint commitment with explicit work-in-progress limits and a pull system. Work is visualised as it actually flows, each column has a limit, and a new item is pulled only when capacity frees up rather than pushed in at a planning event. What follows is that planning becomes continuous and forecasting becomes statistical: instead of "what fits in two weeks", the question is "what is our cycle time distribution and how much is in the system". Kanban is also explicitly evolutionary — it starts with the process you have rather than replacing it — which makes it easier to adopt and easier to adopt in name only. It suits interrupt-heavy work, support, platform and operations teams far better than a sprint commitment does.

Show me a Kanban board with WIP limits.

The limits are the mechanism; a board without them is a status display.

flowchart LR
    A[Options<br/>no limit, ordered] --> B[Analysis<br/>WIP limit 2]
    B --> C[Build<br/>WIP limit 3]
    C --> D[Review<br/>WIP limit 2]
    D --> E[Verify<br/>WIP limit 2]
    E --> F[Done<br/>throughput counted here]
    D -.->|blocked or rejected<br/>pull back, never push forward| C

The first thing to read is that the limits are per column and small relative to the team. Six developers with a build limit of three is deliberate: it forces pairing or helping rather than everyone holding their own item, and it means the team's default response to a blockage is to converge on it rather than to start something new. A limit set to the number of people is not a limit, it is a description.

The review column is where the answer earns its marks. Review and verify are the columns that queue in almost every real team, and they are the columns nobody limits because they feel like waiting rather than working. When review hits its limit of two, nobody may pull from build — which is uncomfortable, is the point, and converts an invisible queue into a visible stoppage that the team must resolve by reviewing.

The dashed edge is the policy most boards lack. Work that fails review moves backwards; it is never pushed forward with a note. Allowing forward movement under pressure is how a verify column fills with items that are not actually done and how cycle time measurements stop meaning anything.

The Options column intentionally has no limit and no start date, because nothing in it has been committed to. The commitment point is the boundary between Options and Analysis, and that boundary is where the lead-time clock starts — a detail that decides which number you are reporting and is worth naming before someone asks why your cycle time looks so short.

Why does limiting work in progress make a team faster?

Because throughput is bounded by the bottleneck, and extra concurrent work adds queueing rather than output. Starting a sixth item when five are in flight does not create capacity; it splits attention, multiplies context switching, and — most importantly — extends the wait time of everything already in the system, so everything finishes later and nothing finishes sooner. Limits also make the bottleneck visible: when a column is full, the constraint is displayed rather than hidden in individual busyness, and the team's only legal move is to help clear it. The cultural difficulty is that a limited board looks idle to a manager watching utilisation, which is the conversation to prepare for. High utilisation and short wait times are mathematically incompatible, and the team is optimising for the second because that is what a customer experiences.

Show me Little's law applied to a real WIP number.

The law is three variables and one of the few pieces of arithmetic that changes a team's behaviour immediately.

Little's law, in the form a delivery team uses:

    average cycle time  =  average WIP  /  average throughput

MEASURED, Checkout team, last 8 weeks
    average items in progress            18
    average items completed per week      6

    cycle time = 18 / 6 = 3 weeks

    So an item pulled today finishes, on average, in 3 weeks - and the
    team was telling stakeholders "about a week, it is only small".

CUT WIP, CHANGE NOTHING ELSE
    WIP 18 -> 6          cycle time = 6 / 6  = 1 week
    WIP 18 -> 9          cycle time = 9 / 6  = 1.5 weeks

    Throughput did not improve. Predictability improved threefold, and
    that is usually the thing being asked for.

THE COUNTER-INTUITIVE PART
    Cutting WIP in half does not halve output. Output was already
    capped by the bottleneck. The 18 items were not being worked on,
    they were waiting - and every one of them was ageing.

CHECK IT AGAINST FLOW EFFICIENCY
    touch time on a typical item   ~2 days
    cycle time                     15 working days
    flow efficiency                2 / 15 = 13%

    87% of elapsed time is queueing. No amount of working harder
    addresses 87%, which is why "the team needs to be faster" is
    almost always the wrong diagnosis.

ASSUMPTIONS, because an interviewer will ask
    - averages over a period where the system is roughly stable
    - items enter and leave at similar rates, nothing abandoned
    - consistent definition of when the clock starts and stops

The first block is the whole argument. A team can report an honest per-item estimate of a few days and deliver in three weeks, with nobody lying, because the estimate describes touch time and the stakeholder heard elapsed time. Little's law is what makes that gap arithmetic rather than an argument about diligence.

The flow efficiency figure is the number to bring to a management conversation about speed. Thirteen per cent efficiency means the intervention with any leverage is in the queues — review latency, environment availability, waiting on another team — and that hiring, overtime and exhortation all address the thirteen per cent. It reframes "go faster" as "stop starting", which is a change the team can actually make unilaterally.

The point that throughput does not fall when WIP is cut is the one that meets resistance, and it is worth being precise: throughput is set by the constraint, so reducing WIP below the point where people are genuinely idle does reduce it. The practical rule is to cut until someone has nothing to pull, then look at why — the answer is almost always a full downstream column, which has just identified the bottleneck for free.

Stating the assumptions matters because the law is exact only for a stable system with consistent boundaries. Teams that abandon a third of their items, or that start the clock at different points for different work types, will get a number that does not reconcile with observed dates, and knowing why is the difference between using the model and quoting it.

What is the difference between cycle time and lead time?

Cycle time measures from when the team started work to when it was done; lead time measures from when the customer asked to when they got it, so it includes the whole period the request sat in the backlog. The distinction matters because they answer different people's questions: cycle time tells the team how its process is performing, and lead time tells the customer what to expect. Reporting cycle time to a stakeholder who asked a lead time question is the most common sleight of hand in delivery reporting, and it is usually accidental — the team improves cycle time from twelve days to five while lead time stays at three months because the queue in front of the board is untouched. The precondition for either number meaning anything is a written, agreed definition of where the clock starts.

What is a blocker policy, and why does a board need one?

It is the team's pre-agreed answer to what happens when an item cannot progress: how it is marked, who owns chasing it, how long it may sit before it escalates, and whether the team may pull a replacement. Without one, blocked items are invisibly parked, the WIP limit is quietly evaded because a blocked item "does not count", and cycle time inflates with no explanation. A workable policy is explicit: blocked items keep their slot and count against the limit, the blockage and its cause are written on the card, anything blocked over two days is raised at the daily and goes on an impediment log with an owner, and the team's first priority each morning is unblocking rather than starting. The log matters more than the flag, because repeated blockers of the same kind are the material for the structural fix.

When should a team drop sprints entirely?

When the sprint commitment has become a fiction the whole team maintains. The clearest signal is unplanned work consistently above roughly a third of capacity — support teams, platform teams, anything on-call — because at that level any two-week plan is invalidated in the first three days and the ceremony cost buys nothing. A second signal is that the team already deploys continuously and the sprint boundary has no relationship to any release, so the timebox is a reporting artefact. What to keep when the sprint goes is the part that mattered: a regular cadence for the retrospective and for stakeholder review, WIP limits in place of the commitment, and throughput-based forecasting in place of velocity. Dropping sprints without any of those is not Kanban, it is an unmanaged queue with a board.

Metrics that mean something

Show me a cycle time distribution read at the 85th percentile.

An average cycle time hides the behaviour that stakeholders actually experience, and the percentile is the number to quote instead.

Cycle time in working days, last 58 completed items, Checkout team.

  days  items
   1    ###                        3
   2    ########                   8
   3    ###########               11
   4    #########                  9
   5    ######                     6
   6    ####                       4
   7    ###                        3
   8    ##                         2
   9    ##                         2
  10    #                          1
  12    ##                         2
  14    #                          1
  18    #                          1
  22    #                          1  <- waited 9 days on the data team
  26    #                          1
  31    #                          1
  38    #                          1  <- blocked, then re-scoped twice
  45    #                          1  <- forgotten in review for 5 weeks

  mean          7.3 days     425 item-days / 58 items. Matches
                             almost no actual item
  median  p50   4   days     "half our work takes 4 days or less"
          p85  12   days     "85% of our work is done within 12 days"
          p95  31   days
          max  45   days

WHAT TO SAY TO A STAKEHOLDER
  "If we start it this week, there is an 85% chance it is done within
   12 working days. Half the time it is 4. About one item in ten
   takes longer than a fortnight, and when that happens it is almost
   always waiting on someone outside this team."

WHAT THE SHAPE TELLS YOU
  - right-skewed with a long tail. Normal for knowledge work. Never
    summarise it with a mean.
  - the tail above 18 days is 5 items. Four of them were blocked on
    an external dependency or sat in review. None was "hard work".
  - so the improvement with leverage is the tail, not the median.
    Halving the median saves 2 days per item; removing the review
    queue removes 20-plus days from the worst items.

The reason to quote p85 rather than the mean is that it is a promise you can keep. A mean of 7.3 days is exceeded by roughly a quarter of items, so a stakeholder planning against it is disappointed regularly; a stated eighty-five per cent service level is both honest about uncertainty and specific enough to plan against. It also converts every future conversation from a debate about estimates into an agreement about a confidence level.

The tail is where the money is, and this is the insight teams miss because improvement effort naturally goes to the common case. Five items over eighteen days contribute more total delay than the entire left half of the distribution, and every one of them was delayed by queueing or a dependency rather than by the difficulty of the work. That points the retrospective at the review column and the data-team dependency, not at the team's pace.

The item forgotten in review for five weeks is worth calling out as a class rather than an anecdote. Nothing on a board tells you an item has stopped ageing well, so the operational fix is an ageing policy — flag anything in a column beyond the p85, discuss it at the daily — which catches the tail while it is still recoverable instead of at the retrospective.

The honest caveat to volunteer is sample dependence. Sixty items is enough for a p85 and thin for a p95, and if the work mix changed halfway through the period the distribution is two distributions overlaid. Splitting by work type — feature, bug, support — usually reveals that and gives three usable service levels instead of one misleading one.

Show me a cumulative flow diagram read for what it reveals.

A cumulative flow diagram is the one chart that shows queueing rather than progress, and reading it is a skill an interviewer can test in a minute.

Cumulative item count by day. Each band is a state. Bands stack, so the
vertical thickness of a band is the WIP in that state on that day.

items
 90 |                                                    ,-DONE
    |                                              ,---''
 75 |                                        ,---''
    |                                  ,---''
 60 |                            ,---''
    |                      ,---''  VERIFY  ####################
 45 |                ,---''        REVIEW  ############################
    |          ,---''                              <- band widening
 30 |     ,---''      BUILD  ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
    |  ,-''
 15 | ,'   ANALYSIS  ------------------------------------------
    |,'
  0 +----------------------------------------------------------
     day 1        day 10        day 20        day 30      day 40

WHAT EACH FEATURE MEANS
  slope of the DONE line       throughput. Flattening = delivery stalled
  vertical thickness of a band WIP in that state today
  horizontal distance between  approximate cycle time for work at that
  the top and bottom curves    point in the period
  a band widening over time    a queue forming. Work enters faster than
                               it leaves that state
  a band with a flat top edge  nothing left that state at all. A total
                               stoppage, usually a blocked dependency
  all bands parallel           a stable system. This is the target shape

WHAT THIS ONE SHOWS
  1  the REVIEW band widens steadily from day 18. Review is the
     bottleneck, and it is getting worse, not oscillating.
  2  the DONE slope shallows after day 25 even though ANALYSIS and
     BUILD keep climbing. The team is still starting work at full rate
     while finishing less - the classic response to a bottleneck.
  3  by day 40 there is more work in REVIEW and VERIFY than was
     completed in the whole period. Cycle time has roughly doubled and
     nobody has noticed, because everyone is busy.

WHAT TO DO
  a WIP limit on REVIEW, so building stops when review is full. That
  feels like slowing down and is the only thing that shortens cycle
  time. Adding people to BUILD makes this chart worse.

The diagnostic value is entirely in the band thicknesses, which is what makes it different from a burndown. A burndown shows a single remaining-work line and would report this period as broadly on track until late; the cumulative flow diagram shows on day eighteen that a queue is forming, three weeks before the delivery line bends.

The second observation is the behavioural pattern the chart exists to expose. When a downstream state is constrained, teams respond by starting more work, because starting is available and finishing is not. That widens the queue, lengthens cycle time for everything in the system, and produces a dashboard where every individual is busy and throughput is falling.

The intervention is counter-intuitive enough to be worth stating explicitly: the fix is upstream of the problem. Limiting review capacity forces the team to review rather than build, which reduces WIP, which shortens cycle time by Little's law. Adding capacity to build — the state that looks busiest — increases arrival rate at the bottleneck and makes every number worse.

The shape to describe as healthy is parallel bands of roughly constant thickness with a steady done slope. Saying that gives an interviewer the reference point, and it is the thing to look for first: any band that is systematically thickening is a queue, and every queue is cycle time being spent on nothing.

Why does a burndown chart flatter a team?

Because it plots remaining estimate rather than completed work, so effort spent on items that are not finished still moves the line down. A sprint where eight stories are all eighty per cent done shows a healthy burndown and delivers nothing, and the drop that reveals it arrives on the last day when the estimates that were never going to close are finally admitted. Two further distortions are worth naming: scope added mid-sprint is invisible unless the chart tracks it separately, so the line can descend through a sprint whose scope grew by half; and the ideal diagonal implies work should complete linearly, which no real system does. A burnup chart showing completed work and total scope as two separate lines fixes most of this, and a cumulative flow diagram or a cycle time scatterplot answers the underlying question far better.

Show me the DORA metrics and how each is gamed.

Four metrics quoted as a set precisely because each is individually gameable.

metric                what it measures         how it gets gamed
--------------------  ----------------------  --------------------------
Deployment            releases reaching       count staging deploys;
frequency             production per unit     split one change across
                      of time                 six deploys; deploy empty
                                              changes on a schedule

Lead time for         first commit to         start the clock at merge,
change                running in production   hiding two weeks of review
                                              queue; measure only
                                              hotfixes, which are fast

Change failure        share of deploys        reclassify incidents as
rate                  causing degradation     "planned maintenance"; fix
                      or needing remediation  forward so no rollback is
                                              recorded; raise the bar
                                              for what counts

Failed deployment     time from impact to     start the clock at
recovery time         service restored        acknowledgement rather
                                              than at customer impact;
                                              close at mitigation and
                                              track the real fix apart

THE PAIRING THAT RESISTS GAMING
  throughput  = deployment frequency + lead time
  stability   = change failure rate + recovery time
  Movement in one pair at the cost of the other is a trade, not an
  improvement. The honest claim is one pair better, other pair flat.

WHAT DORA DOES NOT MEASURE
  whether the right thing was built. Whether anyone uses it. Team
  health. Design quality. A team can be elite on all four while
  shipping features nobody wanted, which costs more than a slow
  pipeline does.

Every entry in the right-hand column is achieved by moving a definition rather than by lying, which is what makes it hard to detect and easy to do without intending to. The defence is procedural: definitions written down, owned outside the team being measured, and instrumented from deployment records and incident timestamps rather than self-reported into a spreadsheet.

The lead time row is the one worth pressing in any interview, because the gaming is nearly universal and usually unconscious. Most tooling measures merge to deploy, which is the part that is already automated, while the fortnight an item spent waiting for review — the actual constraint — sits outside the measurement. The useful question is never "what is our lead time" but "which part of it is queueing".

The pairing is the structural point. Deployment frequency alone rewards recklessness and change failure rate alone rewards shipping nothing, so the set only works as a set. Attaching a target or a bonus to any single one produces the right-hand column immediately, which is Goodhart's law operating on a quarterly cycle.

The final block is what separates a candidate who has used these from one who has read the report. DORA measures the delivery pipeline and says nothing about product value, so a team should hold it alongside a product outcome measure. The benchmark clusters are also self-reported survey data, which makes comparing your team to the elite cohort much weaker than comparing your team to itself last quarter.

What is throughput, and why is it a better forecasting input than velocity?

Throughput is the number of items completed per unit of time, counted rather than estimated, and that single property is why it forecasts better. It requires no estimation step, so it carries no estimation error and cannot be inflated by recalibrating a scale; it is denominated in a unit — items — that means the same thing to everyone; and because it is measured after the fact it already includes every tax on the team, from leave to on-call to a slow review queue. Combined with random resampling it produces a probability distribution over dates rather than a single number. The caveat to volunteer is that it assumes item sizes stay roughly comparable, so a team that starts slicing much smaller will show rising throughput with unchanged delivery, which is why the forecast is sanity-checked against actual dates rather than trusted blindly.

Which metrics should never be reported upward as a team score?

Anything the team itself defines or can trivially adjust, which means velocity, story points completed, estimate accuracy, sprint commitment attainment and individual ticket counts. Each of these is gameable at zero cost and undetectably from outside — estimate higher, commit to less, split more finely — so the predictable outcome of measuring them is a better number and unchanged delivery, plus the loss of the diagnostic value they had when they were nobody's target. The measures that survive being reported are the ones defined outside the team and instrumented from systems: throughput of items, cycle time percentiles, change failure rate, and whatever product outcome the work was meant to move. The sentence worth having ready is that a metric can be a diagnostic or a target but not both, and choosing target converts an honest signal into a managed one within about two sprints.

What would you measure to tell whether an agile transformation worked?

Outcomes rather than adoption, and the distinction is the whole answer. Adoption metrics — teams trained, ceremonies held, tools rolled out, framework certifications — measure that the transformation happened to people, not that anything improved, and they are what transformation programmes report because they are easy and always green. The defensible set is lead time from request to production, deployment frequency, change failure rate, the proportion of delivered work that moved a stated product metric, and something about the people, such as voluntary attrition and whether engineers say they have autonomy over how work is done. The honest addition is a baseline: most transformations cannot demonstrate improvement because nothing was measured beforehand, and saying that you would spend the first six weeks establishing one is a stronger answer than any list of target values.

Scaling and dependencies

Show me a dependency map and the reduction rather than the management.

Every multi-team programme produces a dependency register; the useful move is deleting rows from it rather than tracking them.

flowchart TD
    A[Checkout team<br/>owns the sprint goal] --> B[Platform team<br/>needs a new API field]
    A --> C[Data team<br/>needs an event schema change]
    A --> D[Identity team<br/>needs a new token scope]
    B --> E[Reduce - checkout raises the PR<br/>platform reviews only]
    C --> F[Reduce - schema owned as a contract<br/>consumer-driven test, no ticket]
    D --> G[Delete - move the check into checkout<br/>no cross-team work at all]

The three reductions on the right are three distinct patterns, and knowing them by shape is more useful than any tracking mechanism. The first is inner sourcing: the dependent team writes the change and the owning team reviews it, which converts a scheduling problem — waiting for another team's next planning cycle — into a code review. It requires the owning team to accept contributions, which is a real organisational commitment and the reason it usually fails.

The second is turning a coordination point into a contract. If the event schema is versioned and covered by consumer-driven tests, the producing team can change it safely and the consuming team learns of a breaking change from a failing build rather than from a meeting. The dependency still exists technically and has stopped consuming anybody's planning capacity, which is the part that was expensive.

The third is deletion, and it is the one candidates never reach for. If the identity check can live inside checkout with the data it already has, the dependency was an artefact of where a boundary was drawn rather than a fact about the problem. Asking "why does this need to cross a team boundary at all" reframes the question from scheduling to design, and the answer is often that the team boundary is in the wrong place.

What none of these is: a dependency register reviewed weekly by a coordinator. That is management, it makes the dependency visible and no cheaper, and its steady state is a programme where the constraint is the calendar. The measure that keeps this honest is the proportion of a team's sprint items requiring another team, because a team above roughly a quarter cannot forecast its own work.

Why is a dependency to be removed rather than tracked?

Because tracking reduces surprise and does nothing about cost. A cross-team dependency imposes a queue — the other team's planning cycle, priorities and capacity — so the item's cycle time is set by someone else's calendar, and a register makes that visible without shortening it. Worse, a well-run register creates the impression the problem is handled, which removes the pressure that would have produced a structural fix. The reductions worth naming are giving the dependent team the ability to make the change themselves, replacing the coordination with a versioned contract and consumer-driven tests, moving the capability into the team that needs it, or redrawing the boundary so the work is internal. The signal that this has gone wrong is a team whose sprint plan is mostly conditional on other teams, which cannot be fixed by better tracking at any level of diligence.

What is SAFe, and what is the honest criticism of it?

SAFe is a prescriptive framework for coordinating many teams, built around release trains — five to twelve teams, or fifty to a hundred and twenty-five people, on a shared cadence — with a planning event every eight to twelve weeks and defined roles above team level. The fair case is that large organisations genuinely need a mechanism for cross-team alignment, that big-room planning surfaces dependencies which would otherwise be found late, and that it supplies a common vocabulary. The honest criticism is that it reintroduces much of what agile was reacting against: a planning increment is a quarterly batch, alignment comes through added roles and layers, and the prescription is heavy enough that adopting it is usually a reorganisation. The critique with most force is that it makes an existing structure work slightly better while leaving the dependencies that caused the coordination problem untouched, and autonomy is what pays. Hold neither endorsement nor contempt: it is a reasonable tool for an organisation that cannot yet reduce its dependencies, and the goal is needing less of it each year.

What did the Spotify model actually say, and how is it misused?

It was a description of how one company worked at one point in 2012, published as two articles and a video, describing squads, tribes, chapters and guilds — and its authors have since said repeatedly that it was never a model and that Spotify itself moved on. The misuse is adopting the org chart and skipping everything that made it function: high autonomy backed by genuine alignment on outcomes, teams owning their services in production end to end, minimal handoffs, and a strong engineering culture that tolerated local variation in process. What organisations typically implement is the renaming — teams become squads, departments become tribes — with the same approval chains, the same shared platform team as a bottleneck, and no change in who decides anything. The general lesson worth stating is that copying another company's structure copies the artefact of their context, and the interesting question is always which constraint their structure was solving and whether you have it.

What is LeSS, and when would you choose it over SAFe?

LeSS is Scrum scaled by descaling: one product backlog, one product owner, one definition of done and one shared sprint across several teams, with as little added structure as possible. Coordination happens through teams talking directly and through joint refinement rather than through new roles, and the framework deliberately adds almost nothing to Scrum — its guidance is largely about what to remove, including component teams, handoffs and coordinator roles. You would choose it over SAFe for a small number of teams on one genuine product where the organisation is willing to change structure, because it preserves team autonomy and a single ordered backlog. You would not choose it where there are dozens of teams across several products with fixed funding cycles and hard external dependencies, since it assumes an appetite for organisational change that most enterprises adopting a scaling framework do not have.

How do you run a distributed team without a daily standup ritual?

By moving the coordination into writing and reserving synchronous time for the things that genuinely need it. In practice that is an asynchronous written update against the sprint goal — what is blocked, what changed, what I need from someone — posted in one place before a stated time, plus a board kept current enough to be the actual source of truth rather than a reconstruction. The synchronous meeting then becomes shorter, less frequent and about specific problems rather than round-the-room status, and it is scheduled in whatever overlap exists rather than at one region's convenience. Two disciplines make it work: decisions are written down where someone waking up six hours later can find them, and the overlap window is protected for conversation rather than filled with meetings. The failure mode to name is the team that adopts async updates and keeps the meeting, doubling the cost.

What breaks first when a single team becomes four?

The single ordered backlog and the shared understanding, in that order. One team had one list and one product owner making trade-offs; four teams typically acquire four lists, four local priorities and no mechanism for deciding between them, so the organisation's actual priority becomes whoever escalates. Immediately behind that comes ownership: work that was internal now crosses boundaries, and the way those boundaries are drawn decides everything. Teams organised around components generate a dependency for every feature, while teams organised around end-to-end customer capabilities can mostly deliver alone, which is the same argument as vertical against horizontal slicing applied to the org chart. The third casualty is the informal communication that made the single team work, and the recognisable symptom is that integration becomes an event with its own schedule rather than a continuous property.

Interview traps

Why is "we follow the Scrum Guide" a weak answer?

Because it answers a question about judgement with a citation. Every interesting delivery question — a team that cannot finish anything, a product owner who will not prioritise, a stakeholder demanding a fixed date — has a framework-shaped answer that is technically correct and does not describe what you would actually do on Monday. Quoting the Guide also invites the follow-up that ends the discussion badly: what did you do when the organisation would not give you that, because most organisations will not. The stronger shape is to name what the practice is for, say what you did when the ideal was unavailable, and be explicit about the trade you accepted. Interviewers hiring a scrum master are usually hiring someone to work inside a compromised system, so fluency in the compromises is the signal, and purity reads as inexperience.

Why is "I would remove the impediment" not an answer about impediments?

Because it restates the accountability instead of describing the work, and the work is where the difficulty lives. Real impediments are mostly outside the team's control — another team's priorities, an environment nobody owns, a product owner with no mandate, an approval process measured in weeks — and none of them is removed by a scrum master noticing them. A usable answer names a specific impediment, describes the sequence attempted, and says what happened: the data it took to make the cost visible, who was persuaded and who was not, what was escalated and to whom, and what was worked around in the meantime. It should also be honest that some impediments are not removed but managed down, because admitting that is more credible than a story where every blockage yields. The detail interviewers listen for is whether an impediment log exists and what came off it.

What do interviewers hear when a candidate defends the ceremonies?

That they may be the person who made agile unpopular. Engineers on an interview panel have usually experienced the events as overhead, and a candidate whose answer to "the team finds the standup useless" is that the standup is mandatory has signalled which side of that they are on. The move that works is to concede the observation and then diagnose it: the standup is useless in a recognisable way — status to a manager, round the room, nobody listening — and that specific failure has a specific fix. Defending the purpose while conceding the implementation is the position that survives, because it demonstrates you can tell the mechanism from the ritual. The candidates who fail this are usually the ones who have only ever run the events as specified and have never had to justify one to a hostile, competent engineer.

How should you answer "how do you handle a product owner who will not prioritise"?

By diagnosing before acting, because the four causes need different responses. If they lack the mandate — a proxy with stakeholders above them who override — the work is organisational and the answer is getting the real decision-maker into the review or making the escalation explicit. If they lack information, supply it: cost of delay, a forecast showing that the current list is nine months of work, and the dates that fall out of the ordering they refuse to make. If they are avoiding the conflict of telling a stakeholder no, the answer is to make the trade visible so the refusal is arithmetic rather than personal. And if they are simply absent, that is an impediment to escalate rather than a gap for the team to fill. The move interviewers listen for is making the consequence of no decision visible, because an unordered backlog means the team is prioritising, silently and without the context to do it well.

Why is "velocity went up forty per cent" the wrong success story?

Because it is the one claim in delivery that can be produced without changing anything, and an experienced interviewer knows it. Estimating the same work higher raises velocity immediately, costs nothing and is undetectable from outside the team, so the number is evidence about the team's calibration rather than its output. Even taken at face value it is the wrong quantity: nobody outside engineering wants more points, they want features sooner and fewer failures. The version of the story that lands uses measures the team cannot define — lead time from request to production fell from eleven weeks to three, spillover from six items a sprint to one, change failure rate from a quarter to under a tenth — and attaches them to a mechanism, such as a WIP limit on review or vertical slicing. Outcome plus mechanism is a claim that can be interrogated, which is why it is believed.

What is an interviewer testing when they ask how you would fix a team that keeps missing its sprint commitment?

Whether you diagnose before prescribing, and whether you can consider that the commitment itself is the defect. A strong answer refuses to treat the miss as the problem and asks what the data says: how much of each sprint was unplanned work, what the cycle time tail looks like, whether the same large items spill repeatedly, where items sit on the board and for how long, and whether the sprint has a goal at all. It then separates the likely causes, because they need opposite responses — stories horizontally sliced, an invisible queue in review, unplanned work above a third of capacity, planning to full capacity, cross-team dependencies, or a commitment set by someone other than the developers — picks one, changes it, and measures. It is also willing to say that for some teams the fix is to stop committing to a sprint scope and run flow with WIP limits and a throughput forecast. A weak answer goes straight to a remedy, or reaches for the team's discipline and motivation, which is the diagnosis interviewers are listening for and the one that is almost never correct.