Six months into a scaled agile rollout the sponsor asks whether it is working. What do you measure, and what do you refuse to report?
Report the interval a customer experiences and the quality of what arrives, measured end to end across the programme rather than aggregated from teams. Refuse adoption counts, aggregated velocity and anything that a team can improve by reclassifying its own work.
What the interviewer is scoring
- Whether the candidate asks what the rollout was supposed to fix before choosing any measure
- That the difference between adoption and outcome is stated as a property of the measure rather than as a preference
- Does the candidate know that team-level cycle times cannot simply be aggregated into a programme number
- Whether at least one measure is refused out loud, with the mechanism by which it would be gamed
- That the absence of a baseline is confronted rather than worked around
Answer
Short answer
Report the interval a customer experiences and the quality of what arrives, measured end to end across the programme rather than aggregated from teams. Refuse adoption counts, aggregated velocity and anything that a team can improve by reclassifying its own work.
Flow metrics are the cleanest way to tell whether the rollout changed delivery rather than vocabulary. Useful flow metrics show whether work moves faster, waits less, ages less, and crosses team boundaries with fewer surprises. Weak flow metrics count ceremonies, training completion, or framework adoption. Strong flow metrics connect the transformation to lead time, throughput, dependency ageing, escaped work, and whether teams can make decisions closer to the customer. Flow metrics should make the rollout accountable to delivery reality.
Establish what was supposed to change
Before any measure, get the sponsor to restate the problem in the terms they used when they funded this. There are only a few candidates — releases took too long, quality was poor enough to consume the roadmap, nobody could tell what was coming, the same work was being built twice — and each implies a different measure. A rollout funded to shorten time to market is answered with an interval; one funded to make delivery predictable is answered with the width of a forecast, and those two can move in opposite directions, because predictability is often bought with slack that makes the median slower.
If the sponsor cannot restate it, that is the finding, and it is worth saying at the six-month review rather than assembling a dashboard that answers nobody's question. A programme with no stated objective produces exactly the reporting everyone complains about: a page of counts, all green, that nobody can act on.
Adoption measures are always green, and that is what is wrong with them
Teams trained, ceremonies held, boards migrated, certifications awarded, planning events run, percentage of teams "on the framework" — these are the numbers a transformation office reports, and every one of them rises monotonically because the programme controls both the activity and the counting. They are measures of the rollout happening to people. Their reliable property is that they cannot go down, which means they carry no information: a measure that only moves one way tells you nothing you did not know when you funded it.
The trap in the phrasing is that these measures are not merely useless. They are actively harmful, because a sponsor holding a green adoption dashboard has a defensible reason not to look at the delivery numbers, and the programme survives on that basis for a second year. If you report anything of this kind, report it as cost — this is what we spent — rather than as progress.
Which measures survive contact with a scaled programme
The measures worth reporting have one property in common: the team being measured cannot improve them alone by changing how it describes its own work.
| Measure | Definition that survives | How it fails at programme level |
|---|---|---|
| Lead time to customer | Request accepted to a user having it, per item, as a distribution | Meaningless if the clock starts at team pickup rather than at the request |
| Deployment frequency | Deploys reaching production, counted per service | Fine, and says nothing about whether anything valuable shipped |
| Change failure rate | Share of changes needing a fix, hotfix or rollback | Falls when teams stop classifying incidents honestly |
| Time to restore | Detection to service restored | Robust, and only visible if incidents are recorded consistently |
| Flow distribution | Share of delivered items that were features, defects, debt or risk | The strongest programme measure, and nobody collects it |
| Forecast width | The gap between an eighty-fifth-percentile date and a fiftieth | Needs eight or more comparable intervals of history |
| Proportion of work needing another team | Items in a period with a cross-team dependency | The measure that tells you whether the scaling is still needed |
| Percentage of items that moved a product metric | Delivered items with a stated, measured outcome | Requires product measurement the programme may not own |
The last two rows are where a coaching answer separates itself from a delivery-reporting answer. The share of work requiring another team is the only measure on the list that reports on the reason the scaling framework was bought. If the framework is working as a transitional mechanism, that number falls year on year and the coordination apparatus can be reduced. If it is flat after two years, the programme has become permanent overhead, and saying so with a number is the most useful thing a coach does in the role.
Flow distribution deserves its place too, because it answers the question a sponsor actually has and no other measure does. A programme delivering the same throughput as last year, where defects and unplanned risk work have gone from a fifth of items to half, has got materially worse while every speed measure held. Nothing but a classification of the work reveals that.
The aggregation error that invalidates most programme dashboards
Two arithmetic mistakes recur, and an interviewer will look for whether you know them.
The first is averaging percentiles. If four teams each report an eighty-fifth-percentile cycle time of eight days, the programme's eighty-fifth percentile is not eight days — it is worse, often much worse, because the programme's distribution is the union of four distributions and its tail is built from every team's tail. Percentiles must be recomputed from the pooled items, never averaged from team figures.
The second is summing throughput. Four teams completing ten items a sprint do not give a programme completing forty customer-visible things, because a share of each team's items exist only to unblock another team's item. Counting at the team level counts the coordination as output. The corrective is to measure the interval and the count at the boundary a customer touches, and to treat internal handoffs as cost rather than delivery.
That is also why measures get gamed rather than met, and the mechanism is worth drawing because it is the same shape every time.
flowchart TD
A[Measure published<br/>as a programme target] --> B{Can a team move it<br/>without changing<br/>what a customer gets}
B -->|yes| C[It will be moved<br/>within two quarters]
B -->|no| D[Survives as a measure]
C --> E[Smaller slices<br/>later clock start<br/>work reclassified]
E --> F[Number improves<br/>lead time to customer<br/>does not]
D --> G[Paired with a quality measure<br/>so speed cannot be borrowed]The branch to look at is the second box, because it is a test you can apply before publishing anything rather than a diagnosis afterwards. Cycle time fails it — moving the clock start from request to pickup halves it with no change in delivery. Deployment frequency fails it partially, since a deploy of nothing counts. Lead time measured from the customer's request passes, which is precisely why it is the one nobody wants to instrument.
Say out loud what you will not report
Refusing is part of the answer, and naming the mechanism rather than the objection is what makes the refusal credible. Aggregated velocity across teams: the unit differs per team, so the sum is arithmetic on incompatible scales, and publishing it makes estimate inflation the cheapest available response. Team-against-team comparisons of any flow measure without the work mix beside them: a platform team and a checkout team are not comparable and the comparison teaches both to game. Individual throughput or points: it produces smaller stories and less helping, both immediately and undetectably. Commitment reliability as a percentage: it is a measure of how conservatively a team commits, and it improves fastest when a team stops attempting anything uncertain.
Offer a substitute in each case rather than only a refusal, because a sponsor who wants a single number will get one from somewhere, and it will be worse than the one you declined to give.
The defect that quietly decides the whole review
Nobody measured anything before the rollout. Six months in, this is the position roughly every transformation is actually in, and the two available responses separate a serious practitioner from an evangelist. The weak response is to report the current numbers as though they demonstrate improvement, which nobody can check and everybody suspects. The strong one is to reconstruct a baseline from evidence that already exists — deployment history in the pipeline, ticket timestamps from before the change, incident records, release notes — say precisely how far back it is trustworthy, and present the honest verdict, which is often that the first six months bought visibility rather than speed.
That answer is also the one that survives the follow-up, because it treats the review as an experiment with an inconclusive result rather than a case to be made. A coach who cannot say "we do not yet know, here is what would tell us by March" has no way to recommend stopping, and a transformation that cannot be stopped is not being evaluated at all.
Any measure a team can improve by relabelling its own work will be improved that way, so the test before publishing one is whether a customer would notice the improvement.
© 2026 Preptima. Originally published at preptima.com.
Likely follow-ups
- You have no baseline because nobody measured before the rollout. What do you do at the six-month review?
- How would you report a measure that has got worse for a defensible reason?
- Which single measure would you give a team for its own use and never show upward?
- What would convince you to recommend abandoning the rollout rather than adjusting it?
Related questions
- You have six teams building one product and they keep blocking each other. How do you coordinate the dependencies, and what do you make of frameworks like SAFe?hardAlso on scaling6 min
- A team tells you the scaled agile rollout is process for its own sake and wants no part of it. What do you do?hardAlso on scaling7 min
- Three teams building one product are on different cadences and each waits on the others. Would you align their sprints?hardAlso on scaling7 min
- Leadership wants DORA metrics on a dashboard for every team. How do you use them well, and how would each one be gamed?hardAlso on goodharts-law5 min