Skip to content
QSWEQB

Product management fundamentals

The answers a PM loop is built from: what the role actually decides, evidence that survives scrutiny, prioritisation defended rather than recited, metrics read correctly, and the traps in a product sense round. Sixty items, thirteen worked through with a table, a calculation or a diagram.

60 questions

Go deeper on Product Management

What the role decides

What does a product manager actually own, and what do they not?

You own the problem, its priority and the definition of success; you do not own the solution, the design or the estimate. That split is the whole job. Owning the problem means you are accountable for whether the thing built was worth building, which is why "engineering delivered what I asked for and nobody used it" is your failure and not theirs. Not owning the solution means the design and the architecture belong to the people whose craft that is, and a PM who arrives with wireframes and a schema has spent their leverage on the part of the work someone else does better. The authority question follows from this: you rarely have formal power over anyone, so the decisions you make stick only to the extent that the reasoning behind them is visible and holds up when argued with.

What is the difference between a product manager, a project manager and a product owner?

A product manager decides what to build and why; a project manager decides how the agreed work gets coordinated and lands on time; a product owner is a role defined inside Scrum, accountable for the backlog and its ordering. The distinction that matters in an interview is scope of question. A project manager optimises delivery of a decision already taken, so their failure mode is a perfectly run project that shipped the wrong thing. A product owner as Scrum defines it is a subset of product management — backlog, acceptance, sprint-level trade-offs — with no necessary claim on strategy, pricing or market. Many organisations use "product owner" as a job title for a backlog administrator with no discovery mandate, and if you are interviewing for one it is worth asking directly which version they mean, because the two jobs are not the same career.

What is the difference between an outcome and an output?

An output is a thing you shipped; an outcome is a change in behaviour that followed. "We launched the referral programme" is an output, "invited-user signups went from four to eleven per cent of new accounts" is an outcome. The distinction is not vocabulary policing — it changes what you commit to. A team committed to outputs is done when the feature is live and has no reason to look at it again, whereas a team committed to an outcome has to keep working until the number moves or has to conclude, publicly, that the approach was wrong. That second option is the expensive one and the reason most roadmaps are lists of outputs. The honest caveat to give an interviewer is that some work genuinely has no measurable outcome in a quarter — a compliance deadline, a platform migration — and pretending otherwise produces invented metrics.

Show me the four risks applied to one feature.

Every feature carries four independent risks, and the point of naming them separately is that each is tested by a different activity.

Feature: in-app invoice scanning. Photograph an invoice, get a draft
expense record with supplier, date and amount pre-filled.

risk         the question              how it fails             cheapest test
-----------  -----------------------  -----------------------  ------------------
VALUE        will they choose it       they keep keying by      6 interviews about
             over what they do now     hand because the photo   last month's actual
                                       step feels slower        entry, not the idea

                                                                fake-door: put the
                                                                button in, count
                                                                taps, no feature
                                                                behind it

USABILITY    can they succeed at it    OCR returns a wrong      5-person prototype
             unaided                   amount and the user      test with their own
                                       does not notice, so      real invoices, not
                                       the fix is invisible     clean samples

FEASIBILITY  can we build it with      accuracy on crumpled     2-day spike: run 200
             what we have, in the      phone photos of thermal  real invoices through
             time we have              receipts is 60%, not     the candidate vendor
                                       the 95% on the demo      and score by field

VIABILITY    does it work for the      OCR cost per account     unit-economics sheet
             business - legal,         exceeds the margin on    plus one call each
             margin, support, sales    the tier that wants it;  with support, legal
                                       support cannot triage    and finance
                                       a wrong amount

Verdict here: feasibility and viability are the binding risks, not value.
Test those first, in that order, before any design work.

The reason to separate them is that teams reliably test the wrong one. Value gets tested obsessively because interviews and surveys are cheap and pleasant, while feasibility is deferred to the sprint where it becomes a problem, and viability is often not tested at all until support and finance object after launch.

The ordering rule is to attack the risk that would kill the feature soonest and costs least to resolve. Here the vendor accuracy spike takes two days and can end the project, so doing it before a single screen is designed saves the design work rather than validating it. A team that runs the interviews first learns that people want this — which everyone already believed — and learns nothing about whether it can exist.

Usability is the one candidates collapse into value. They are different: people can want a thing and still fail to use it, and the failure mode here is particularly nasty because a wrong amount that looks confident is worse than an empty field. That specific failure is what a prototype test with real, messy invoices surfaces and a clean demo hides.

Viability is the risk PMs are uniquely responsible for, because nobody else in the room is accountable for margin, legal exposure or the support load. Naming it unprompted is one of the fastest ways to sound like a PM rather than a requirements author.

Who decides when the PM and the engineering lead disagree?

Whoever owns the axis being disagreed about, and most of these disputes go away once that is said out loud. If the disagreement is about whether the problem is worth solving now, it is yours. If it is about how the solution is built, the sequencing of technical work, or what the estimate is, it is theirs. The genuinely contested cases are trade-offs across the boundary — six weeks of hardening against two features, or a shortcut that ships in April and costs a year — and those resolve by making the cost explicit in the other's currency rather than by seniority. If it still deadlocks, escalate with a written recommendation and both positions, and accept the answer visibly. A PM who wins by escalation habitually finds that estimates start arriving padded.

What does a PM owe engineering, and what does engineering owe a PM?

You owe them a prioritised problem rather than a specification, the reasoning behind the priority so they can make a hundred small decisions without asking, a decision when they need one instead of a deferral, and enough stability that the plan is not re-cut weekly. They owe you honest forecasts with the uncertainty attached, early warning of a slip rather than a heroic recovery, options with costs instead of a flat refusal, and the technical constraints translated into consequences you can weigh. Both failure modes are recognisable. A PM who specifies solutions removes the judgement they hired engineers for and gets compliance instead of ideas. An engineering lead who hides risk until it is unrecoverable removes your ability to manage the expectations you set on their behalf.

Why is "CEO of the product" a misleading description?

Because it describes authority you do not have and omits the skill the job actually requires. A chief executive can hire, fire, allocate budget and overrule; a PM typically controls none of those and gets outcomes by assembling evidence, framing trade-offs and persuading people who do not report to them. The phrase also sets a bad posture for new PMs, who read it as licence to direct and then find that engineers and designers route around them. The more accurate framing is that you are accountable for the outcome and responsible for the decision quality, with influence as the only instrument — which is why the standard interview probe is a time you changed a decision without authority, and why "I escalated" is a weak answer to it.

Discovery and evidence

What is discovery actually for?

Reducing the cost of being wrong before you spend a quarter being wrong. Discovery is not a phase that produces requirements; it is a continuous activity that produces evidence about which of your beliefs are load-bearing and which are guesses. The practical output is not a document but a changed decision: a feature descoped, a problem reframed, a segment ruled out. That is the test to apply to your own discovery — if six weeks of research changed nothing about what you were already going to build, it was theatre. The other thing to say is what discovery cannot do. It cannot tell you what people will pay for by asking them, and it cannot validate an idea, only fail to falsify it, which is why the strong version is framed as risks to disprove rather than an idea to confirm.

Show me a discovery interview script with the leading questions rewritten.

Most interview scripts collect agreement rather than information, and the fix is mechanical.

Context: accounting SaaS. We suspect manual invoice entry is painful and
are considering scanning. Nine interviews, small-business tier.

LEADING                              REWRITTEN
-----------------------------------  -----------------------------------------
"Would you use a feature that        "Walk me through the last invoice you
scanned invoices with your phone?"   entered. What did you do first?"
  Asks about a hypothetical self.      Asks about an event that happened. The
  Everyone says yes to a free           story contains the workflow, the
  feature. Answer carries no            workarounds and the cost, none of
  information.                          which they would have thought to say.

"Isn't manual data entry             "How much time did you spend on invoice
frustrating?"                        entry last month, and how do you know?"
  Supplies the emotion and the         Makes them produce a number and its
  answer. They will agree to be         source. "No idea" is itself a finding:
  polite.                               unmeasured pain is rarely prioritised
                                        pain.

"How important is accuracy to        "Tell me about the last time an expense
you, on a scale of one to five?"     figure turned out to be wrong. What
  Nobody says accuracy is             happened next?"
  unimportant. The scale gives          Gets the consequence, which is what
  false precision to a                  tells you whether accuracy is a
  meaningless answer.                    preference or a compliance problem.

"Would you pay ten pounds a          "What do you spend money on today to
month for this?"                     make this less painful? A bookkeeper,
  Stated willingness to pay is        another tool, overtime?"
  close to worthless. People            Revealed spending is evidence. Existing
  are generous with imaginary            spend is also the price anchor.
  money.

"We're building this next            "If this existed and worked perfectly,
quarter - any concerns?"             what would you stop doing?"
  Announces a decision and asks       Forces a trade-off. If the answer is
  for permission. Concerns will        "nothing, I'd use both", the pain is
  not be raised.                       not big enough to displace a habit.

CLOSING MOVE, every interview
  "Who else does this in your company, and can I talk to them?"
  A referral is behavioural evidence that the conversation was worth
  their time. Refusals cluster in the segments where the pain is weak.

The single transformation running through the column is from the future conditional to the past indicative. People are unreliable narrators of what they will do and fairly reliable reporters of what they did, so every question that can be pointed at a specific past event should be.

Numbers you ask them to produce are more useful than numbers you offer them to rate. A five-point scale about importance produces a distribution centred just above the middle for every question ever asked, whereas "how long did it take and how do you know" separates the people who have measured the problem from the ones who have merely complained about it.

The willingness-to-pay rewrite is the one interviewers probe hardest. Stated price tolerance is systematically inflated and unanchored, while existing spend on workarounds — a bookkeeper, a second tool, weekend hours — is real money already leaving the account and tells you both that the pain is priced and roughly at what.

The referral close is the cheapest validity check in the script. It converts politeness into a cost: someone who will spend their colleague's time on this believes the problem is real, and a run of nine interviews with two referrals is a much weaker signal than the transcripts will make it feel.

What is the mom test principle?

Ask questions your mother could not distort by loving you — that is, questions about her life rather than about your idea. In practice it means three rules: talk about their past behaviour instead of your product, ask about specifics rather than generalities, and never mention the solution until you have the problem story, because the moment the idea is on the table the conversation becomes about your feelings. The reason it works is that compliments, hypothetical enthusiasm and feature requests are all forms of politeness, and none of them constitute evidence. The practical tell that you have run a bad interview is that you left feeling encouraged and cannot name a single fact you learned that would have changed a decision.

When do you need quantitative evidence and when will qualitative do?

Qualitative tells you why and what to build; quantitative tells you how many and whether it worked. They answer different questions and substituting one for the other is the common error in both directions. Eight interviews cannot tell you that thirty per cent of users have a problem, and an analytics funnel cannot tell you why they abandon at step three — it can only tell you that they do, precisely and reliably. The sequence that works is qualitative to generate hypotheses and find the mechanism, quantitative to size it and to measure the change. The sharpest thing to say in an interview is what each cannot do: interviews have no denominator, and event data has no explanation, so a claim resting on one alone should name which half is missing.

What is survivorship bias in product feedback?

It is the systematic absence of the people whose opinion would have changed your mind. Your feedback channels — support tickets, NPS responses, the customer advisory board, the users who reply to your email — are populated almost entirely by people who stayed, and disproportionately by the most engaged of those. So you optimise for power users, harden the workflows they already tolerate, and never hear from the eighty per cent who tried the product for nine minutes and left. The corrective is to go looking for the missing population deliberately: churned accounts, activation failures, trials that never converted, and the accounts your sales team lost. It is more work and less pleasant, which is exactly why the data you already have feels sufficient.

What is jobs to be done, and what does it replace?

It is the framing that people buy progress in a situation rather than a product, so the unit of analysis is the job — "get an accurate expense figure before the quarter closes" — rather than the customer's demographics. What it replaces is segmentation by attribute, which routinely groups people who behave differently and separates people who behave the same. The practical payoff is that it widens your view of competition: the competitor for invoice scanning is not another app, it is the bookkeeper, the spreadsheet and doing nothing. The honest limitation is that jobs are inferred rather than observed, so a badly written job statement is just a feature description with aspirational language wrapped around it, and it is easy to produce one that cannot be wrong.

How do personas get misused?

They stop being a summary of research and become a cast of invented characters with stock photographs, at which point they are used to settle arguments rather than inform them. Three specific misuses recur. Personas built from assumptions rather than interviews, which launder opinion into apparent evidence. Personas defined by demographics — "Marketing Mary, 34, likes yoga" — when nothing in the demographic predicts the behaviour you care about. And personas used as a permanent artefact, printed and pinned up, long after the product and the customer base moved. The version that survives contact with a team is small, behavioural, dated, and cites the research it came from, so that anyone can check whether the claim being made in the person's name is actually in the data.

Show me an opportunity sizing calculation done aloud.

Interviewers want the assumptions narrated, not the answer, so the arithmetic is deliberately simple and every line is contestable.

Question: how big is invoice scanning for our accounting SaaS?

BASE
  paying accounts today                              12,000   known
  small-business tier, who key invoices by hand         55%   billing data
  addressable accounts                                6,600

WHO HAS THE PAIN
  in 9 interviews in that tier, 6 named manual
  entry as a top-three monthly annoyance
  taking that as roughly 60% - a weak number
  from a tiny sample, flagged as such               x  60%
  accounts with the pain                              3,960

ATTACH
  comparable paid add-ons attached 10-20% of the
  accounts with the pain in year one. Take the
  middle, 15%                                       x  15%
  paying accounts, year one                              594

REVENUE
  594 x GBP 15/month x 12                          GBP 106,920
  round to                                         GBP 107k ARR

COST TO SERVE
  OCR vendor GBP 0.02/page, est. 200 pages per
  account per month = GBP 4/account/month
  594 x 4 x 12                                     GBP  28,512
  gross contribution                               GBP  78k

COST TO BUILD
  2 engineers + 0.5 design, 4 months, loaded
  ~GBP 11k/person-month                            GBP 110k

READ  payback in year two, on a 15% attach rate that is itself the
      softest number in the model. Halve the attach rate to 8% and it
      still pays back, just later: 3,960 x 8% = 317 accounts, at
      GBP 11 net per account per month that is GBP 41.8k a year, so
      the GBP 110k build is recovered during year three.

SENSITIVITY  move each input by 1% and see what happens to the
      GBP 78k gross contribution:
        price        +1.36%   594 x GBP 0.15 x 12 = GBP 1,069
        attach rate  +1.00%   contribution is linear in attach
        OCR cost     -0.36%   594 x GBP 0.04 x 12 = GBP 285
      Price is the most leveraged input, because the net margin per
      account is GBP 11 on a GBP 15 price, so a price move lands
      almost entirely on the bottom line.

WHAT I WOULD DO  spend the next money on attach rate anyway. It is
      not the most leveraged input, it is the least evidenced one -
      six of nine interviews and a borrowed 10-20% range - whereas
      price is a number we set and can revisit at any time. A
      fake-door test measuring intent to attach costs a fortnight and
      moves the widest error bar in the model.

The purpose of the exercise is not the £107k. It is the identification, in the final lines, of which input to go and reduce the uncertainty on — and an interviewer is listening for whether you find it, because that is what tells you what to do next.

Notice that the answer is not simply "the input with the biggest coefficient". Price has the larger sensitivity here, 1.36 against 1.00, and it is still the wrong thing to research, because we already know the price and can change it on a Tuesday. Attach rate has a smaller coefficient and a far wider error bar, and you can only narrow that bar by going and measuring something. Leverage tells you where the answer moves; evidence tells you where you are guessing, and the next experiment belongs to the second question.

Every questionable number is labelled at the point of use. "Six of nine" is stated as a weak basis rather than smoothed into sixty per cent and then treated as a fact three lines later, which is how sizing models acquire false confidence. Saying "this is the number I am least sure of" is a strength in this format, not a hedge.

The two cost lines are what separate sizing from revenue fantasy. A model that stops at £107k of ARR has not answered the question that was asked, because the decision is whether to spend the quarter, and the quarter costs £110k. Volunteer cost to serve in particular: per-unit vendor costs against a fixed subscription price are the standard way a feature is revenue-positive and margin-negative.

The sensitivity block is also the answer to the follow-up about what you would do with more time, and it is worth showing the arithmetic rather than asserting a ranking. Each figure is one line: a one per cent price rise adds fifteen pence a month across 594 accounts, which is £1,069 a year against a £78k contribution, so 1.36 per cent. Doing that for three inputs takes a minute and stops you claiming, as models of this kind routinely do, that the answer is insensitive to something it is in fact most sensitive to.

The pay-back line deserves the same treatment. "It never pays back at 8%" is the sort of sentence that sounds appropriately cautious and is simply false on the model's own numbers, and an interviewer who checks it will conclude that you did not. Halving the softest input and reporting the honest consequence — a year later, not never — is both more useful and harder to argue with.

Strategy and positioning

What is a product strategy, and how is it different from a roadmap and a backlog?

A strategy is a choice about where to compete and how you intend to win there, stated in a way that rules things out. A roadmap is the sequence of problems you intend to work on next in service of that choice. A backlog is the inventory of work items currently under consideration. They differ by what they constrain: a strategy tells you which customers and problems you are declining, a roadmap tells you order, a backlog tells you nothing at all about direction. The test of a real strategy is that it is falsifiable and costly — if reversing every sentence would still be an acceptable thing for the company to say, you have written a mission statement. Most organisations that believe they lack a roadmap actually lack the strategy the roadmap would have been derived from, which is why the roadmap keeps being relitigated.

How would an interviewer expect you to use TAM, SAM and SOM?

TAM is total demand for the category if you had all of it, SAM is the part your model and geography can actually serve, SOM is the share you could realistically capture in a defined period. What an interviewer is testing is whether you can do the bottom-up version and whether you know what the numbers are for. Top-down from an analyst report — "a $40bn market, one per cent of which is $400m" — is the answer that fails, because the one per cent is unargued. Bottom-up is number of accounts times price times attach rate, built from figures you can defend. And the purpose is comparative: TAM tells you whether a market can support the ambition, SOM tells you what to plan against, and quoting TAM in a business case where SOM is the relevant figure is the most common misuse.

What is positioning, and how is it different from messaging?

Positioning is the decision about what category you are in, who you are for and what you are better at; messaging is the words used to express that decision. Positioning comes first and is expensive to change, because it determines who compares you to whom and therefore which features you are expected to have. The useful discipline is to state it as a sentence with real alternatives in it: for this segment, who have this job, we are the option that does this better than that named competitor, and here is why that is credible. Teams that skip it end up writing messaging for a position they never chose, which is why their landing page claims to be simpler than the enterprise tools and more powerful than the simple ones, and converts nobody.

How do you find a differentiator that holds?

Look for something a competitor could copy but will not, because copying it would cost them something they are unwilling to lose. Anything a well-resourced rival can add in a quarter is a feature, not a differentiator — which rules out most feature lists. What holds tends to be structural: a data asset that accumulates with use, a distribution channel they cannot access, a business model that would cannibalise their existing revenue, or a deliberate narrowness that would force them to serve their main segment worse. The interview version of this question is often "why won't Google do this", and the weak answer is that you will move faster. The strong one names the conflict that makes the copy unattractive rather than difficult.

Why does a dated roadmap cause trouble, and what do you publish instead?

Because every date on it is read as a commitment by someone whose plans then depend on it, while the dates furthest out are the ones you know least about. Sales quotes them, marketing books campaigns, customers write them into renewal conversations, and when the third quarter moves — as it must, since discovery has not happened yet — the cost is not a schedule change but credibility. What to publish instead is a horizon view: now, next and later, where "now" carries real dates because it is in build and scoped, "next" names the problem with the solution still open, and "later" is direction with explicit permission to change. The essential part is stating the confidence attached to each band, because the trouble comes from uniform presentation of non-uniform certainty.

Show me now, next and later against a dated Gantt.

The same three problems, published two ways, and the difference is entirely in what the reader is entitled to expect.

flowchart TD
    A[Three problems<br/>one plan, two ways to publish it] --> B[Dated plan<br/>SSO 14 Mar, audit log 30 Jun,<br/>mobile access 30 Sep]
    A --> C[NOW<br/>enterprise login pain<br/>scoped and in build<br/>date committed at 85 per cent]
    C --> D[NEXT<br/>audit and compliance gap<br/>problem agreed, shape open<br/>quarter named, no date]
    D --> E[LATER<br/>field access<br/>direction only<br/>may be dropped entirely]
    B --> F[Read as three promises<br/>two of which are guesses<br/>and one will be quoted to a customer]
    E --> G[Read as intent with<br/>one promise inside it<br/>and permission to learn]

The two branches contain identical intentions. The only difference is that the right-hand path tells the reader how much weight each band can carry, and that single addition is what stops the plan being used as a contract.

The dated version fails in a specific sequence worth describing, because interviewers have lived through it. The March date is roughly right, since that work is scoped. The September date was invented in January to fill the row. Sales quotes September to close a deal in April, discovery in July reveals the mobile problem was actually an offline problem, and the choice becomes shipping the wrong thing on time or breaking a commitment somebody else made on your behalf.

The now band still has dates, and it is important to say so. Refusing to date anything is the over-correction that makes a roadmap useless to everyone who has to plan around you, and it reads as evasion. Committed, scoped work in build gets a date with a confidence attached; nothing else does.

The later band needs its disposability stated in words. If it is not explicit that items here may be dropped, the band functions as a queue with an implied guarantee, and the difference between "later" and "never" is precisely the thing you are trying to communicate. The health check on the whole artefact is whether anything has ever moved backwards out of next — if not, it is a dated plan with the dates hidden.

Show me a build, buy or partner decision as a table.

The decision is rarely about cost, and framing it as a build-versus-licence price comparison is how it gets made badly.

Need: OCR and field extraction for invoice scanning.

                    BUILD                BUY / licence        PARTNER
                    in-house model       vendor API           white-label a
                                                              scanning product
------------------  -------------------  -------------------  ------------------
Time to first       6-9 months           2-3 weeks            4-6 weeks, mostly
customer value                                                commercial

Year-one cost       GBP 250k+ of         GBP 28k at forecast  revenue share,
                    engineering,         volume, superlinear  ~20% of the
                    ongoing              with growth          add-on's revenue

Is it our           No. Nobody buys us   No                   No
differentiator      for OCR quality

Switching cost      n/a                  moderate: an         high: their brand
later               low, we own it       abstraction layer     and UX are in our
                                          keeps it 2 weeks     product

Key risk            we are not an ML     price rises or the   they get acquired
                    team; accuracy       vendor is acquired;  or deprioritise
                    plateaus below       data residency for   us; we cannot fix
                    the bar              EU accounts          their bugs

Data exposure       none                 invoices leave our   invoices leave,
                                          estate - needs DPA   plus they own the
                                          and EU region        customer relation

DECISION  Buy, behind our own interface. Build is a category error: this is
          table stakes, not differentiation, and we have no ML capability.
          Partner gives away the customer relationship on a workflow we
          intend to own. Revisit build only if volume makes per-page cost
          exceed roughly GBP 150k a year, which is about 5x the GBP 28k
          forecast in the row above.

The row that decides it is "is it our differentiator". Build is defensible when a capability is the reason customers choose you, and almost never otherwise, because the true cost of building is not the first version but the maintenance, on-call and improvement of a component you will always be worse at than a specialist.

The switching-cost row is the one candidates omit, and it is what makes the decision reversible. Buying behind your own abstraction converts a strategic commitment into a two-week replacement, which is why the decision line says "behind our own interface" rather than just "buy" — that clause is the entire risk mitigation for the vendor-acquisition risk named above it.

Partner is dismissed for a reason unrelated to cost, which is the point of having the row about the customer relationship. Putting someone else's product inside a workflow you intend to own means you cannot fix its bugs, cannot control its roadmap, and have taught your customer to attribute the value elsewhere.

Naming the condition that would reverse the decision is what turns this from an opinion into a decision record. "Revisit at five times volume" gives the choice a review trigger, and an interviewer will usually ask for exactly that, because a build-buy answer with no reversal condition suggests you think the answer is permanent.

Prioritisation

What does a prioritisation framework actually give you?

A structured argument, not an answer. The output of any scoring model is only as good as the inputs, all of which are estimates you supplied, so a framework cannot tell you what to build — it can only make your reasoning visible enough to be challenged, and force comparison on consistent axes rather than on whoever spoke most recently. That is genuinely valuable: most bad prioritisation is not bad arithmetic but incomparable arguments, where one feature is justified by revenue, another by strategy and a third by a customer's volume. The failure to avoid is treating the score as authoritative. If the model ranks something first and everyone in the room knows it is wrong, the useful move is to find which input encodes the disagreement, not to override the score quietly.

Show me a RICE table scored for four candidate features.

Four candidates, nine engineer-weeks of quarter, and the argument is in the columns rather than the ranking.

Scale used, stated because RICE is meaningless without it:
  reach     accounts touched per quarter, from analytics
  impact    3 massive, 2 high, 1 medium, 0.5 low, 0.25 minimal
  conf      100% we have data, 80% some evidence, 50% a guess
  effort    engineer-weeks, from the tech lead, not from me

feature             reach   impact  conf   effort    RICE   rank
                  accts/qtr           %   eng-wks
-----------------  --------  ------  ----  -------  ------  ----
Bulk CSV import         220     1.0   100        2   110.0     1
Report scheduling       160     0.5    80        1    64.0     2
Mobile access           900     1.0    50       20    22.5     3
SSO for enterprise       40     2.0    80        6    10.7     4

Arithmetic:  220 x 1.0 x 1.00 / 2  = 110.0
             160 x 0.5 x 0.80 / 1  =  64.0
             900 x 1.0 x 0.50 / 20 =  22.5
              40 x 2.0 x 0.80 / 6  =  10.7

What I would do: take import and scheduling - 3 weeks, ranked 1 and 2,
no argument. Then take SSO with the remaining 6 weeks despite it
ranking last, and defer mobile.

THE WEAKNESS IN THE SCORE
  Reach counts accounts, and accounts are not equal. The 40 accounts
  wanting SSO are the enterprise tier: GBP 240k of ARR in renewal in
  November, and two of them have named SSO as a blocker in writing.
  The 220 accounts wanting CSV import average GBP 180 a year.
  Weighted by revenue at risk, SSO ranks first and the model cannot
  see it, because RICE has no term for value per unit of reach.

The ranking is not the answer, and saying so is the answer. The score is a tool for surfacing the disagreement, and here it surfaces one immediately: three of the four rows are consistent with intuition and the fourth is not, which is a signal to interrogate an input rather than to accept the output.

The input at fault is reach, and the flaw is structural rather than a scoring mistake. RICE multiplies reach by impact per user, so it systematically favours broad, shallow wins over narrow, deep ones and is blind to revenue concentration. Any product with an enterprise tier will find that RICE ranks its enterprise work last, every quarter, until someone notices.

Confidence is the second place the model hides things. Mobile access is scored at fifty per cent, which is an admission that the whole row is a guess, and a fifty sitting in a table next to a hundred looks like a number rather than the sentence it actually is: we do not know if this matters. The right response to a low confidence is usually a cheap test rather than a ranking.

Effort belongs to the tech lead, and taking it from anywhere else is where these tables go wrong quietly. A PM who supplies their own effort estimates has a model that produces whatever they already wanted, and the twenty-week mobile figure is exactly the sort of number a PM would have optimistically halved.

Where does RICE break down?

In four places worth naming, because the follow-up question is always this one. Reach and impact multiply, so it prefers wide shallow features and undervalues work that is critical to a small, valuable segment. It has no term for dependencies or sequencing, so it will happily rank a feature above the platform work it requires. It cannot represent time-sensitivity — a compliance deadline and a nice-to-have score identically if their other inputs match — which is what cost of delay exists to fix. And precision is illusory: two significant figures computed from a guessed confidence and an estimated impact invite comparisons the inputs cannot support. Used as a conversation structure it is genuinely useful; used as a decision procedure it produces confident nonsense.

What is the Kano model useful for?

Distinguishing kinds of value rather than amounts, which is the thing linear scoring cannot do. It sorts features into must-haves, whose absence causes dissatisfaction but whose presence earns nothing, performance attributes where more is proportionally better, and delighters that produce satisfaction out of proportion to their cost. The practical use is deciding where investment stops: you cannot win on a must-have, you can only fail to lose, so pouring a quarter into making your login the best in the category is wasted. It also predicts decay — today's delighter becomes tomorrow's must-have as competitors copy it, which is why the analysis has to be redone rather than treated as a fixed classification of your product.

What is cost of delay, and when is it the right lens?

It is the value lost per unit of time that a thing is not shipped, which reframes prioritisation from "what is most valuable" to "what is most expensive to postpone". It is the right lens whenever value is time-dependent: a regulatory deadline where the cost is a step function, a seasonal window that closes, a competitor's launch that erodes an advantage, or a fixed cost you keep paying until a migration completes. Dividing it by duration gives the weighted-shortest- job-first ordering, which is the defensible way to justify doing a small urgent thing before a large valuable one. The difficulty is honest quantification, and the usual failure is asserting a large cost of delay for something whose real deadline is that a stakeholder would like it sooner.

What is the difference between prioritisation and sequencing?

Prioritisation ranks by value; sequencing decides the order you actually do things in, and the two differ because of dependencies, capacity shape and risk. The highest-priority item may not be first: it may depend on platform work, need a designer who is busy for three weeks, or be the item whose main risk is best retired by a spike done alongside something else. Sequencing also accounts for what a team can hold at once — four parallel high-priority items usually finish later than the same four done two at a time. The confusion causes a specific and common argument, where a stakeholder concludes their item was deprioritised when it was in fact sequenced second for a reason nobody explained, and the fix is to publish the ordering with its constraints attached.

How do you handle an executive's pet feature?

Convert it from an instruction into a hypothesis without making it a contest. First find the actual belief underneath it, because a pet feature is usually a solution attached to a real observation — a customer conversation, a competitor's demo, a board question — and the observation is often worth acting on even when the solution is not. Then price it publicly against what it displaces, in the currency they care about, and propose the cheapest test that would settle it. If they still want it after that, ship the smallest honest version, instrument it, and agree beforehand what result would mean stopping. Two things not to do: refuse outright, which converts a disagreement into a status fight you will lose, and comply silently while resenting it, which loses you the credibility to object next time.

How do you prioritise reliability and technical debt against features?

By translating it into the same currency as everything else rather than asking for goodwill. Debt and reliability work compete badly when presented as engineering hygiene and compete well when presented as their consequences: this defect class generated forty per cent of last quarter's support tickets, this fragility is why the estimate on every change in that area is doubled, this incident rate is why two enterprise renewals asked about uptime. Then the trade-off is between named costs rather than between features and virtue. The mechanisms that work in practice are a standing allocation the team controls without asking, and attaching the improvement to the feature that touches the same code. What does not work is a debt backlog, which is a list of things nobody will ever prioritise.

Metrics

Show me a north star metric with the leading indicators beneath it.

A north star is only useful with the tree underneath it, because the metric itself moves too slowly to act on.

flowchart TD
    N[North star<br/>weekly active teams that<br/>ship at least one report] --> A[Activation<br/>new teams reaching a first<br/>report within 7 days]
    N --> B[Engagement<br/>reports per active<br/>team per week]
    N --> C[Retention<br/>week-4 teams still<br/>shipping a report]
    A --> A1[Data source connected<br/>in the first session]
    B --> B1[Template used rather<br/>than a blank start]
    C --> C1[Second seat invited<br/>within the first fortnight]

The choice being made at the top is the substance. "Weekly active teams that ship a report" was picked over registered users, logins and revenue because it is the smallest unit of the customer getting the value the product exists to deliver — and because it is a team-level count, which matches how the product is bought and how it spreads.

Each layer down trades slower and truer for faster and noisier. The north star moves over quarters, the three drivers move over weeks, and the bottom row moves within a session and can therefore be the target of a shipped change. That is the whole purpose of the tree: it gives a team something they can affect on Tuesday that is causally connected to the thing the company cares about.

The bottom row is where a tree earns its keep or becomes decoration, and each item there should be a claim you could be wrong about. "Second seat invited in the first fortnight" asserts that single-user teams churn — a testable belief, and if the correlation turns out to be spurious the branch gets deleted rather than quietly optimised.

What deliberately does not appear is revenue. Revenue is the outcome the north star is a leading indicator of, and putting it at the top produces a metric the team cannot move directly and that improves for reasons — a price rise, a large deal — having nothing to do with the product. The honest caveat to volunteer is that a single metric always under-describes the business, which is what the guardrails exist to cover.

What is the difference between a leading and a lagging indicator?

A lagging indicator reports a result — revenue, churn, quarterly retention — and a leading indicator predicts one early enough to act on. The trade is reliability against timeliness: lagging numbers are unambiguous and arrive too late to change, leading ones arrive in time and may be measuring a correlation that does not hold. You need both, and the mistake is managing exclusively by either. Running on lagging indicators means every course correction is a post-mortem; running on leading ones alone means a team can hit every proxy while the business declines. The practical discipline is to check periodically that your leading indicator still predicts the lagging one, because these relationships decay as the product and the customer base change, and nobody notices for two quarters.

What makes a metric a vanity metric?

That it can only go up, and that no decision changes depending on its value. Total registered users, cumulative downloads, page views and total revenue-to-date all qualify: they are monotonic, they rise even as the product declines, and there is no number at which anybody would do anything differently. The diagnostic to apply is to ask what you would do if the figure halved — if the answer is nothing, or if the figure cannot halve, it is decoration. The useful replacements are almost always rates, ratios or cohort-based: active rather than registered, retention rather than signups, revenue per account rather than total. Naming a vanity metric in your own past work is a strong move in an interview, because the failure is so universal that claiming never to have made it is implausible.

What is activation, and why is it defined badly so often?

Activation is the point at which a new user has experienced enough of the product's value to plausibly come back, and defining it badly is the most consequential metric error a team makes because everything upstream gets optimised towards the wrong finish line. The two bad definitions are the trivial one and the aspirational one. "Completed signup" is trivial and measures nothing about value, so growth work optimises the form rather than the product. "Used nine features in the first month" is aspirational, satisfied by almost nobody, and useless as a target. The defensible way to set it is empirical: find the early behaviour that best separates the cohort that retained from the cohort that did not, then check the relationship is plausibly causal rather than merely a marker of people who were already keen.

Show me a retention curve that flattens against one that does not.

Two products, the same acquisition, and only one of them has a business.

Day-N retention: % of a signup cohort active on that day

           d1    d3    d7   d14   d30   d60   d90
        ------------------------------------------
  A       44    31    24    21    20    19    19
  B       46    30    19    12     7     4     2

  A   44 |*
         |  31
         |     24  21  20  19  19      flattens at ~19%
         |     ------------------
         |
  B   46 |*
         |  30
         |     19
         |         12   7   4   2      decays to zero
         |
         +------------------------------------------
          d1  d3   d7  d14 d30 d60 d90

The flat section of curve A is the finding, and its height is the number that matters. A stable nineteen per cent means roughly a fifth of every cohort found something they keep coming back for, so the product has a retained core and acquisition compounds: users accumulate, and the long-run active total is approximately the acquisition rate multiplied by that floor.

Curve B has no floor, and this is fatal in a way early numbers actively conceal. B's day-one and day-three figures are better than A's, so B looks like the stronger product for the first fortnight, and a team measuring signups and week-one activity would celebrate. Every user B acquires eventually leaves, which means total actives plateau at whatever the marketing spend sustains and the business is a treadmill that stops the moment spend does.

The diagnostic response differs entirely. For A, the work is either raising the floor — which means retention work on the mechanism that made the nineteen per cent stick — or pouring acquisition into a funnel that now demonstrably holds. For B, acquisition spend is destructive, and the only useful work is finding whether any sub-segment flattens: if one does, the product is mispositioned rather than valueless.

The measurement caveat worth volunteering is that "active" has to be defined as the product's core action rather than as opening the app, and that day-N retention and rolling-window retention give different-looking curves for the same product. Also say how long you need: a curve cannot be declared flat over thirty days for a product with a monthly usage rhythm, and a great deal of premature celebration comes from calling a flattening at d30 that resumes falling at d90.

Show me a cohort table and what it says.

The same data as a single active-user line would have looked like growth throughout.

Monthly signup cohorts, % still active in month N after signup

cohort    size    m0    m1    m2    m3    m4    m5
-------  ------  ----  ----  ----  ----  ----  ----
Jan       1,200   100    38    29    26    25    25
Feb       1,450   100    37    28    26    25     -
Mar       2,900   100    26    17    14     -     -
Apr       3,400   100    24    15     -     -     -
May       3,100   100    35     -     -     -     -

Aggregate monthly actives over the same period: 4,100 -> 4,900 ->
6,200 -> 7,400 -> 8,100. Up every month.

Read the table down a column, not across a row. The m1 column goes 38, 37, 26, 24, 35 — a two-month collapse in early retention followed by a recovery — and that is invisible in the aggregate active-user count, which rose every single month because the cohorts were getting larger. This is the specific reason cohort tables exist.

The shape of the collapse tells you where to look. It started in March, affected March and April intake only, and reverted in May, while the January and February cohorts continued along their normal path to a stable twenty-five per cent floor. An event that damages new cohorts but not existing ones is not a product regression — a broken feature would have hurt everyone — so the suspects are acquisition-side: a paid campaign into a poorly matched audience, a change in landing-page promise, or a discount that bought the wrong customers.

The recovery in May is the evidence that settles it, and note how little of the table that claim is allowed to rest on. May has exactly one comparable cell, m1 at 35 per cent, against 38 and 37 for January and February and 26 and 24 for the two damaged months. That single cell is enough: May's cohort is nearly as large as April's and its early retention is back in the pre-March band, which rules out an argument that retention necessarily degrades as you scale acquisition. Something specific happened for two months and stopped, and the next step is to check what changed in the channel mix in March and again in May.

What it is not allowed to rest on is a May m2 figure, because there is not one yet. The diagonal, or right-hand edge, of a cohort table is always the least reliable part, and the newest row is not merely unreliable but absent: a cohort that signed up in May has not lived two months, so any number in that cell was extrapolated by somebody. This is the standard cohort-table error and it is seductive precisely when the incomplete row is the one carrying your conclusion, which is why May's m2 onward is left blank rather than estimated. If the m1 cell had not been enough on its own, the correct answer would have been to wait a month rather than to fill the gap in.

Show me an A/B test result table read correctly.

Two variants, one clear winner, and the interesting reading is of the one that did not win.

Experiment: onboarding checklist. Primary metric: 7-day activation.
Randomised at account level. Pre-registered: MDE 1.5pp absolute,
power 80%, alpha 5%, baseline 17.0% -> 10,200 accounts per arm.
Ran 19 days, which enrolled about 14,500 per arm, so the test is
somewhat over-powered against the pre-registered MDE.

variant     accounts  activated   rate   abs diff   95% CI          z      p
---------   --------  ---------  -----  ---------  -------------  ----  -------
A control     14,610      2,483  17.0%        -          -           -       -
B checklist   14,502      2,741  18.9%   +1.9pp   +1.0 to +2.8   4.24  0.00002
C product     14,388      2,509  17.4%   +0.4pp   -0.4 to +1.3   1.00     0.32
  tour

Guardrails
                            control       B         C
  support tickets/100         3.1        3.0       4.4   <- C worse
  d30 retention              24.1%      24.6%     23.9%
  median time to first
    report                  4.2 days   3.6 days  4.4 days

Decision: ship B to 100%. Discard C. Do not ship B+C together on the
assumption the effects add.

B is a clean result, and the confidence interval is what to quote rather than the p-value: the effect sits between +1.0 and +2.8 points, entirely above zero. The over-reading happens next. The bottom of that interval is +1.0, below the 1.5-point MDE pre-registered as worth shipping — the point estimate clears the bar and the pessimistic end does not. So ship B on the balance of evidence while saying plainly that the effect may be smaller than the threshold. Had 1.5 points been a hard commercial floor rather than a planning figure, the answer would be to keep running until the interval sat wholly above it.

The p-value is worth a sentence only because people quote it wrongly. Here z is 4.24, and a z of 4.24 corresponds to p of about 0.00002, not 0.001 — a factor of roughly forty. Nothing in the decision changes, which is precisely why the error survives in so many read-outs: a p-value is being used as a badge rather than as a quantity, and nobody checks a badge.

C is the row that separates candidates. The correct statement is that no effect was detected, not that there is no effect: the interval runs from -0.4 to +1.3, so a commercially meaningful improvement of a point is entirely consistent with this data. "Non-significant" means the data cannot distinguish the result from zero, and if a +1pp effect would have been worth having, the experiment was underpowered for the question actually being asked.

C's guardrails would have blocked it regardless: support tickets rose from 3.1 to 4.4 per hundred accounts, a real cost landing on another team. That is the case for pre-registering guardrails — without them, a variant that nudges the primary metric and damages something expensive ships, because nobody measured the damage.

The sample-size line is worth defending, because a pre-registration nobody can reproduce is decoration. Detecting 1.5 points on a 17.0% baseline at 5% alpha and 80% power needs about 10,200 accounts per arm; this run enrolled roughly 14,500, which buys power rather than a different conclusion. Quoting a figure that in fact corresponds to 91% power as though it were the 80% one is a small error with a large tell: the design was copied rather than calculated.

The last line is a trap set deliberately. Two variants tested against a shared control tell you nothing about their combination — they may compete for the same attention, and effects on a bounded metric rarely add. Shipping the union of two winners is a distinct experiment presented as a conclusion.

What does statistical significance actually mean, in plain terms?

It means that if there were genuinely no difference between the variants, a result at least this extreme would be unlikely — conventionally under a one-in-twenty chance. That is all. It is not the probability that your variant is better, it is not a measure of how large the effect is, and it says nothing about whether the difference is worth shipping. The two errors that follow are worth naming. A significant result on a tiny effect is common with large samples and frequently not worth the code, so significance and materiality are separate questions. And a non-significant result is not evidence of no effect, only an absence of evidence for one, which is why the confidence interval is the more informative thing to report and the thing a strong candidate reaches for first.

What is the minimum sample, and why is stopping early cheating?

The minimum sample is the number the test needs, calculated before it starts, from your baseline rate, the smallest effect worth detecting, and the power and significance you want. Its purpose is to fix the stopping point in advance. Stopping early is cheating because a running test fluctuates, and if you check repeatedly and stop when it crosses significance, you have optimised over your own noise — the false-positive rate rises well past five per cent, sometimes above thirty with frequent peeking. The result is a portfolio of shipped wins that do not replicate and a metric that never moves in aggregate. The legitimate ways to look early exist and have to be chosen up front: sequential testing, alpha spending, or a pre-declared interim check with an adjusted threshold.

What is a guardrail metric?

A metric you do not intend to improve but refuse to damage, declared before the experiment so it cannot be reinterpreted afterwards. Its job is to catch the displaced cost: a checkout change that raises conversion and doubles refunds, a notification that lifts engagement and raises unsubscribes, an onboarding simplification that improves activation and floods support. The useful set spans functions rather than staying inside your own funnel — latency, error rate, support contacts, churn, revenue per user — because the cost of a local optimisation usually lands on somebody else's team. Two rules make them work: they are chosen before the result is known, and a breach blocks the launch by default rather than starting a negotiation. Without the second, a guardrail is a note.

Delivery and trade-offs

Show me a PRD skeleton with what belongs in each section.

Length is not the problem with most PRDs; the problem is that the sections engineers actually need are the ones missing.

1  PROBLEM AND EVIDENCE                                    half a page
   Who has it, how you know, how often, what it costs them.
   Cite the research: interview count, the funnel step, the ticket
   volume. No solution language anywhere in this section.
   Engineers read this to make the hundred small decisions you will
   not be present for.

2  WHY NOW
   What changed - a deadline, a competitor, a contract, a newly
   available capability. If nothing changed, say so. This is the
   section that survives a reprioritisation argument.

3  SUCCESS AND FAILURE
   The metric, its current value, the target, and the date you will
   read it. Then the number at which you would revert or stop.
   Guardrails listed here, not discovered later.

4  SCOPE
   In scope, as user-visible capability rather than as tasks.
   OUT OF SCOPE, explicitly and generously - this is the highest
   value section in the document and the one most often absent.
   Deferred, with the condition under which it returns.

5  REQUIREMENTS AND RULES
   The behaviour that is not negotiable, especially the unglamorous
   parts: permissions, what happens to existing data, error and
   empty states, limits, currency and locale, audit needs.
   Every rule an engineer would otherwise have to guess or invent.

6  OPEN QUESTIONS, with an owner and a date against each
   The section that makes the document honest. A PRD with no open
   questions is either trivial or lying.

7  DEPENDENCIES AND RISKS
   Other teams, vendors, legal or security review, data migration.
   Named people, not team names.

8  LINKS
   Designs, the technical design doc, the analytics spec, the
   research. The PRD does not contain the solution - it points at it.

NOT IN A PRD: the schema, the API shape, the estimate, the sprint
plan, or the UI copy. Each of those is someone else's document and
putting it here means you will be arguing about the wrong thing.

The problem section carries the whole document. Its purpose is not to justify the work to a reviewer but to give engineers and designers enough context to resolve ambiguity themselves, and a PRD that opens with a solution guarantees a stream of questions to you about every case you failed to anticipate.

Out of scope is the highest-leverage section and the one written last if at all. Scope grows through reasonable local decisions — an edge case here, a settings toggle there — and the only defence is having written down, before the work began, what you were deliberately not doing. It also protects engineers, who otherwise have to litigate each addition without a document behind them.

Section five is where PRDs are actually judged by engineers. Permissions, existing data, error states, limits and locale are the requirements that get discovered in the last week and cause the slip, and they are absent from most PRDs because they are boring to write. A PM who reliably covers them earns latitude everywhere else.

Open questions with owners and dates are what distinguish a working document from a performance. Ambiguity is normal at this stage; unnamed ambiguity is what turns into a two-week stall, and writing "unclear whether we support multi-currency — Priya, by 8 August" costs one line and removes a class of failure.

What makes a PRD useful to engineers rather than ignored?

That it answers questions they would otherwise have to guess at, and stops before telling them how to build it. Concretely: the problem and its evidence, the non-negotiable rules including the unglamorous ones, what is explicitly out of scope, what success will be measured by, and the open questions with owners. What makes it ignored is length without decisions, solution detail that duplicates the technical design, and a version that drifts out of date while the real answers live in a chat thread. The maintenance point is the one most candidates miss: a PRD is useful only while it remains the place where a decision is recorded, so the discipline is to update it when a call is made in a meeting, otherwise the document becomes archaeology within a fortnight.

What is an MVP, properly understood?

It is the smallest thing that produces a real learning about a risk, not a small version of the eventual product. That distinction is the whole point and is almost universally lost. A cut-down version of the full feature answers no question you did not already have an opinion about; a fake door, a concierge process done by hand, a landing page with a real payment step, or a single-segment manual pilot answer specific questions cheaply. The corollary is that an MVP can be embarrassing to build and still correct, and that it may be thrown away by design. What it must never be is a poor-quality version of the real thing shipped to everyone, because that tests your ability to disappoint users rather than any hypothesis, and it contaminates the measurement you were trying to take.

What do you cut when the date will not move?

Scope, in a specific order, and you say which order before the pressure arrives. The first things to go are breadth of cases — one currency instead of four, one integration instead of three — because narrowing the audience keeps the experience whole. Next is automation: a manual back-office step for the first month is invisible to the customer and often reveals that the automation was mis-designed. Next is polish on paths few users take. What you do not cut is anything that makes the released thing untrustworthy: correctness, security, permissions, data integrity, or the ability to tell whether it worked. And you re-verify the deadline first, because a surprising proportion of immovable dates are preferences with a confident tone, and cutting quality for one of those is unforgivable.

Show me a "no" to a customer written out.

The reasoning has to be visible, because a refusal without it is heard as indifference regardless of how politely it is phrased.

To: Head of Finance, Meridian Group - 40 seats, renews in November
Request: a custom approval chain with four configurable levels, tied
         to their internal cost-centre hierarchy.

--------------------------------------------------------------------

Thank you for the detail in Tuesday's call - the cost-centre mapping
in particular was more specific than we usually get, and it changed
how I understand the problem.

I am not going to build the four-level configurable chain, and I want
to be straight with you about why rather than leave it on a roadmap
you will ask me about in three months.

What I heard is that the real problem is that expenses over GBP 5,000
are being approved by people without budget authority, and you find
out at month end. The four-level chain is one solution to that. It is
the most expensive one available to us: configurable approval
hierarchies are roughly a quarter of engineering time, and we have
seen this exact request from two customers out of six hundred, which
means we would be maintaining a subsystem for you specifically and
you would be depending on our attention to it.

What I can do, and what I think solves the actual problem:
  - a spend threshold with a named second approver, per cost centre.
    Three weeks. It covers the over-GBP-5,000 case, which is the one
    with the money in it.
  - an alert to a finance mailbox when anything is approved outside
    the expected chain, so month end holds no surprises. Two weeks.
  - if you need the full hierarchy, our API exposes approval events
    and your team could drive your own chain from it. I will get you
    the documentation and half a day of an engineer's time.

What I will not promise is the configurable version "later". If that
changes it will be because several customers need it, and I will tell
you if it does.

Can we get thirty minutes with your controller to check that the
threshold plus the alert actually closes the month-end gap? If it does
not, I would rather know now than build the wrong three weeks.

The refusal is in the second paragraph, unhedged, and that placement is deliberate. A "no" buried at the end after three paragraphs of appreciation reads as evasion, and the customer will re-read the message trying to work out whether it was actually a no — which they will then resolve by asking again in a month.

The reasoning is given in the company's terms rather than theirs. Two customers out of six hundred, a quarter of engineering, and an ongoing maintenance dependency are facts a reasonable counterpart can evaluate. "It is not on our roadmap" gives them nothing to evaluate and invites escalation to whoever owns the roadmap.

Separating the problem from the requested solution is what makes the alternatives credible. The reframing is stated as something heard from them and offered for correction, not asserted, because if the reframe is wrong the whole response is wrong and it is much cheaper to find that out in this message than after three weeks of work.

Refusing the soft "later" is the hardest part and the part that builds trust. A vague future promise is comfortable to write, defers the argument, and reliably produces a worse conversation at renewal when nothing has happened. Naming the condition under which it would change — several customers, not one — treats them as an adult and leaves the relationship intact.

How do you handle a slipped date?

Report it when the probability moves rather than when the outcome is certain, in writing, with a revised range and the decision you need. The structure is four sentences: what changed, what the new forecast is, what the options are, and what you are asking for. The damage from a slip is rarely the schedule; it is the discovery that your reporting was worthless, which happens when a date is missed the week it was due after a month of green status, and it discounts every forecast you give afterwards. So the discipline is to be visibly wrong sometimes in the optimistic direction. Then handle the downstream commitments yourself — the sales team who quoted it and the customer who planned around it — because that is the part of the cost that is yours rather than engineering's.

How do you work with design without becoming the designer?

Bring the problem, the constraints and the success measure; do not bring a solution. The most common way PMs damage this relationship is arriving with wireframes, which converts a designer into an implementer and loses you the option you would not have thought of. What you legitimately own is the problem definition, the priority, the trade-off when a design is excellent and takes three extra weeks, and the decision when design and engineering disagree about feasibility. Where you push is on evidence — has this been tested with anyone, what happens in the empty and error states, what does this look like for the account with four thousand rows — because those are product questions asked in design's language. Where you do not push is taste, and a PM relitigating visual decisions is a PM whose designer stops showing early work.

Launch and iteration

What are launch tiers, and why do they exist?

They are a pre-agreed classification of how much noise a release gets — from a silent ship to a full campaign — and they exist because launch effort is a scarce shared resource across marketing, sales, support and enablement. A typical scheme runs from tier zero, shipped behind a flag with no announcement, through a changelog entry, an in-product notice and customer email, up to a tier one with press, a webinar and sales enablement. The value is that the conversation about effort happens once, in advance, against criteria, rather than as a negotiation per feature in which whoever lobbies hardest wins. It also protects your customers' attention: a product that announces everything at full volume trains people to ignore announcements, and then the launch that mattered lands silently.

What belongs in a launch readiness check?

The things that are someone else's problem if you forget them. Support briefed and holding written answers to the five likely questions; sales knowing what to say including who it is not for; documentation and in-product help live; the analytics actually verified as recording the events, not merely specified; alerting on the new failure modes; a rollback or flag-off path that has been tested rather than assumed; legal or privacy sign-off where data handling changed; and pricing or packaging configured if it is paid. The one candidates omit is verifying instrumentation before launch — a fortnight of missing events is unrecoverable and means the launch cannot be evaluated at all. The other habit worth naming is deciding beforehand who watches the dashboard on day one and what number would trigger a rollback.

What is a beta actually for?

Learning something specific under conditions you control, not being cautious in general. A useful beta starts with the question it is answering — does this hold up on real data volumes, do people find the entry point unaided, does the support load scale — and selects participants who can answer it, which is usually not your friendliest customers. It also has an end condition and an exit decision agreed before it starts. The failure mode is the beta that runs indefinitely because nobody defined what would end it, quietly becoming a permanently unsupported feature with a small dependent population and no owner. The other trap is treating beta as a quality strategy: shipping something unfinished and labelling it beta transfers the cost to customers and does not make the defects acceptable.

Show me a feature sunset plan.

Removing something is a product decision with a customer-trust cost, and the sequence is what determines whether it is paid.

flowchart TD
    A[Usage and revenue measured<br/>per account, never in aggregate] --> B{Does any account<br/>depend on it materially}
    B -->|no| C[Announce, 60 days,<br/>remove]
    B -->|yes| D[Name the accounts<br/>and their renewal dates]
    D --> E[Build a migration path<br/>or grant a written exception]
    E --> F[Announce with a date<br/>in-product and by email to admins]
    F --> G[Read-only period, then export,<br/>then removal and code deletion]
    C --> G

The first box is where sunsets go wrong. Aggregate usage of two per cent looks like an easy removal and routinely conceals three accounts that use it constantly, one of which is your largest — so the measurement has to be per account, joined to revenue, before any decision is taken.

The branch exists because these are genuinely two different processes. A feature nobody depends on needs an announcement and a waiting period, and treating it with the full migration apparatus wastes a quarter. A feature with material dependency needs a path built for named accounts, timed around their renewals, and the conversation happens before the announcement rather than after it.

The read-only period before deletion is the step most plans skip and the one that prevents the worst outcome. It lets someone who missed every notice discover the change while their data still exists, converting a furious support escalation into an export. Turning something off and deleting its data on the same day is how a minor removal becomes a reference-customer problem.

Code deletion at the end is worth stating explicitly, because a sunset that leaves the feature disabled but present has not delivered the benefit that motivated it. The reason to remove things is the maintenance, security surface and cognitive load they carry, and none of that goes away until the code does.

What are the basics of pricing a product?

Price against the value the customer gets, not the cost you incur or the competitor's number, and choose a value metric that grows with that value — seats, transactions, volume, sites. The value metric is the more consequential decision than the number, because it determines whether revenue expands automatically as accounts get more out of the product or whether every increase requires a negotiation. Beyond that, the standard moves are tiers matched to segments with a deliberate reason for each upgrade, an annual discount for the cash and the retention, and periodic willingness-to-pay research rather than a price set at founding and never revisited. Two things to name unprompted: cost-plus pricing ignores value entirely, and discounting to close deals silently repositions the product for every future negotiation.

How do you decide whether a launch worked?

Against the number you wrote down beforehand, read at the date you said you would read it, and with the honest confounders named. That sounds trivial and is rare: most launch reviews are conducted retrospectively against whichever metric moved, which makes success unfalsifiable. The specific things to hold to are the pre-registered metric and target, a comparison that accounts for the launch bump — adoption in the first fortnight is inflated by curiosity and announcement traffic, so the interesting figure is week four — and the guardrails. Then a decision: double down, iterate, or stop. The last option is the one that has to be genuinely available, because a team that has never killed a launched feature is not evaluating launches, it is documenting them.

What do you do when a launched feature is not used?

Find out which of the four causes it is before doing anything, because the responses are unrelated. Nobody knows it exists, which is a discovery and communication problem and the cheapest to fix. People find it and cannot complete it, which is a usability problem and shows up as an entry-rate that is fine and a completion rate that is not. People complete it and do not return, which means it did not deliver the value and is a value problem. Or the problem was never important enough to displace a habit, which means the discovery was wrong. Funnel data distinguishes the first three in an afternoon; the fourth requires talking to people. The uncomfortable part of the answer is that the fourth case is common and the correct response is usually removal rather than promotion.

Interview traps

What is an interviewer testing with "tell me about a product you would improve"?

Whether you start from a user and a problem or from a feature you fancy. The question is deliberately open, and the structure of your answer is most of the assessment: who the user is and which segment you have chosen, what problem they have and how you would know it is real, what the current experience does badly, two or three candidate solutions rather than one, a reasoned choice between them, and how you would measure whether it worked. A strong answer also names what it is declining to solve and what it would give up. The weak answer is a list of features the product lacks, delivered with confidence and no user in it, which tells the interviewer you would arrive at their company with opinions rather than a method.

Why is "I would talk to users" a weak opening in a product sense round?

Because it is true, universal and answers nothing, and the interviewer is asking what you think — not what you would find out. It reads as a deferral: they have given you a bounded prompt precisely to see how you reason under incomplete information, and outsourcing the reasoning to hypothetical research declines the exercise. The strong version keeps the research and adds the substance around it: state the assumption you are making, say what you believe and why, then name the specific thing you would check and what result would change your answer. That is the same shape as a good design answer — assume, decide, name the falsifier — and it demonstrates the thing being tested, which is judgement in the absence of data rather than the ability to ask for data.

Why is a framework recited without a decision a weak answer?

Because the framework is a tool for reaching a decision, and reciting it substitutes process for judgement. An answer that walks through RICE for four features and then declines to say which to build, or names the four risks without saying which is binding here, has demonstrated vocabulary and nothing else — and vocabulary is table stakes at any level worth interviewing for. Interviewers see this constantly, and the tell is the absence of the word "so". The fix is mechanical: after the structure, commit to an answer, name the assumption it rests on, and say what would change your mind. Committing to something defensible and being argued out of it scores far better than a survey of considerations, because the job consists of making calls with incomplete information and defending them.

What do interviewers hear when you cannot name a metric that would have told you you were wrong?

That you have not actually run a product to a number. Anyone can nominate a success metric after the fact; the harder and more diagnostic question is which value would have made you stop, and a candidate who has genuinely operated this way answers it immediately because they wrote it down at the time. Vagueness here also predicts a specific behaviour on the job — a PM who never sets a failure threshold never kills anything, and their roadmap accumulates features nobody uses because there was never a criterion for removing one. The strong answer states the metric, its baseline, the target, the date it would be read, the number that would trigger a revert, and the guardrail that could veto the whole thing independently.

Why is "we shipped it and users loved it" a weak result?

Because it contains no number, no counterfactual and no cost. "Loved it" is sourced from the people who told you, which is a self-selected sample of the retained and enthusiastic, and it is exactly the evidence a failed feature also generates. A usable result names the metric and its movement, the timeframe, the baseline it moved from, and what else changed at the same time that might have caused it. It is also specific about your contribution, which is what the question is really after: what you decided, what you got wrong on the way, and what the thing cost in engineering time that could have gone elsewhere. Adding the part that did not work is what makes the rest believable, because a story with no cost in it reads as marketing.

What single question most reliably separates candidates in a product management round?

"Tell me about something you decided not to build, and what it cost you." It resists preparation because the answer requires a real decision with a real counterparty, and it exposes whether the candidate has ever exercised the part of the job that is actually scarce. A strong answer names the request and who wanted it — a large customer, an executive, the whole sales team — states the evidence and reasoning that led to the refusal including the numbers, describes what was offered instead, names the cost that was actually paid rather than the one that was feared, and says whether the decision looks right in hindsight, sometimes concluding that it does not. A weak answer describes declining something obviously bad, which is not a decision, or retreats into process: it did not score well on our prioritisation framework. The underlying test is whether the candidate can hold a position against pressure while keeping the relationship, because a PM who says yes to everything has no strategy and a PM who says no without reasoning has no allies.