Your coded case mix has drifted upwards over two quarters. How do you tell better documentation from upcoding?
Only the record can settle it. Rule out mix, population and grouper changes, then run a blind re-code of a stratified sample, report disagreement rate and direction, attribute each disagreement to coder, documentation, template or incentive, and fix the cause rather than the codes.
What the interviewer is scoring
- Does the candidate rule out legitimate causes of drift before treating it as a compliance finding
- Whether the test proposed is a blind independent re-code against the record rather than an analysis of the codes alone
- That disagreements are reported in both directions, so undercoding is found as well as overcoding
- Whether templates, copy-forward text and automated charge triggers are considered as structural causes with no individual author
- Does the answer connect coder and clinician incentives to the metric being measured
Answer
Drift is a question, not an answer
Upcoding means claiming a code that pays more than the record supports. Undercoding is the mirror image, and it is also a defect: it understates the acuity of the population, distorts quality and risk adjustment, and gives away revenue the organisation was entitled to. Framing the problem as "are we overcoding" produces a one-sided investigation and a coding department that learns to code down to stay safe, which is its own compliance exposure.
Start by ruling out the legitimate explanations, because several of them produce exactly the pattern you are looking at.
The service mix may have moved. A new surgical service, a closed clinic, a change in referral patterns or a shift from inpatient to day-case activity all change the case-mix index without anybody coding differently. The population may genuinely be sicker, which two years of deferred care or a demographic shift will do. The coding rules themselves may have changed: classification updates, grouper logic revisions and new guidance all move the distribution on a stated date, and a step change that lines up with a release date is a rules change until proven otherwise. And a documentation improvement programme may be doing precisely what it was funded to do, because clinical notes routinely understate what the clinician knew — an unspecified diagnosis coded specifically after a query is a more accurate claim, not a more aggressive one.
Each of these is checkable against dates and volumes, and each one you can eliminate narrows the question. What you must not do is take the trend as evidence in either direction. A rising case-mix index is not proof of manipulation, and a plausible clinical story is not proof of accuracy.
The only test is the record
The measurement that settles it is a retrospective audit in which an independent, credentialed coder re-codes cases from the documentation without seeing the original codes. Everything else is a screen that tells you where to look.
Design the sample so it answers the question. A stratified random sample gives you an overall accuracy figure you can defend: strata by service line, by coder, by encounter type, by payer, and by the specific code groups that moved. Add a targeted sample on those code groups, and keep the two separate in the reporting, because a targeted sample tells you about a suspicion and a random sample tells you about the population. Mixing them and quoting one accuracy figure is a genuine methodological error that an external auditor will notice.
Report the disagreement rate with direction and cause, not as a single accuracy percentage. The useful categories are: the record does not support the code assigned; the record supported a higher or more specific code the coder did not assign; the record is genuinely ambiguous and two competent coders would differ; and the code is defensible but the sequencing changed the grouping. That last one matters for inpatient work, where which diagnosis is principal and which comorbidities are captured determines the group and therefore the payment, so a case can be wholly correctly coded and still land in the wrong group because of sequencing judgement.
Attribute the cause, because the fix depends on it
A disagreement rate with no cause attached leads to retraining, which is what organisations do when they do not know why something happened.
Coder error is the least interesting cause and the easiest to address. Documentation gaps are the most common and the fix is upstream, with the clinician, not the coder. Structural causes are the ones worth hunting, because they produce upcoding at scale that no individual chose: a note template with a review of systems pre-populated as normal, so every note documents an examination that may not have happened; copy-forward text that carries an acute condition through an admission long after it resolved; an order set that automatically documents severity language; a charge that drops automatically whenever a device is scanned regardless of whether it was used. These generate documentation that looks specific and rich and is not, and they are indefensible in an audit precisely because the record appears to support the code while the underlying clinical activity does not.
Then look at incentives, which is the cause nobody puts in the report. If coders are measured on productivity and revenue per case, or if a service line's performance is judged on case-mix index, the distribution will move and every individual will be able to explain their own decisions. The rule worth stating is that you do not compensate or rank anyone on the coded output; you measure them on agreement with an independent re-code. Whatever you make the metric, you will get.
Clinical documentation improvement queries sit right on this line and are the sharpest control in either direction. A query that describes the clinical findings and asks the clinician to document their assessment is legitimate and produces a better record. A query that names the diagnosis it wants and offers it as the obvious answer manufactures documentation. The distinction is the wording, the wording is retained, and it is discoverable — which makes the query log both your best evidence of a well-run programme and the first thing an auditor reads if it is not.
Benchmarking is a screen with no verdict in it
Comparing your distribution against peers, or one clinician against their colleagues, is cheap and useful for pointing at where to sample. It cannot conclude anything, and treating it as if it can is the trap that catches both sides of this argument.
Being above a peer distribution is expected for a tertiary referral centre, a specialist unit, or any provider whose case mix differs from the comparator group — which is most of them, since the comparator group is rarely constructed to match you. Equally, being within the distribution proves nothing about whether your codes are supported, because a whole peer group can be documenting badly in the same way. Use the comparison to allocate audit effort, then let the record decide. And look at both tails: a coder who never assigns a low-level code and a coder who never assigns a high-level one are both diverging from the record, and only one of them is usually investigated.
What happens when you find it
Confirming that codes were unsupported turns a quality exercise into a compliance one, and the shape of the obligation is worth knowing even where the specifics vary by market. Discovering that you have been paid for claims the record does not support generally starts a clock: the overpayment has to be quantified and returned, and doing so voluntarily is treated very differently from having it found for you. Where a sample shows a systematic error, the question immediately becomes what to do about the population the sample was drawn from, and extrapolating a sampled error rate to that population is a standard method — which is why the sampling methodology needs to be defensible before you start rather than reconstructed afterwards.
The other half is prospective. The corrective action is a change to the template, the order set, the query wording, the charge trigger or the incentive, with a re-audit scheduled to show the rate moved. An audit that produces a training session and no change to the mechanism that caused the error will produce the same finding next year, and the second finding is much harder to characterise as a mistake.
Rising case mix is a hypothesis. The only instrument that tests it is a blind re-code against the documentation, reported in both directions and with each disagreement attributed to a cause you can change — because the fix for a template is not the fix for a coder, and neither is the fix for an incentive.
Likely follow-ups
- Peer benchmarking shows you above the distribution for one code. What does that entitle you to conclude?
- How should a documentation-improvement query be worded, and what makes one indefensible?
- You confirm overpayment on a sample. What obligations does that discovery create, and what do you do about the population you did not sample?
- How would you set up a coder quality programme that does not simply push everyone towards the middle of the distribution?
Related questions
- A GDPR erasure request arrives for a patient whose data sits in an append-only clinical audit log. What do you do?hardAlso on compliance6 min
- Two shipments of the same product were declared under different commodity codes. Why does that matter, and whose problem is it?hardAlso on compliance4 min
- A customer returns three failed units built in different weeks. You have one shift to say what else is affected. How do you bound it?hardSame kind of round: scenario6 min
- The domain expert tells you one thing and the written procedure says another. How do you work out which one the system should follow?hardSame kind of round: scenario5 min
- How would you price a feature that no customer asked for?hardSame kind of round: case-study6 min
- The operations team says there is no rule for this, they just use judgement. How do you model that?hardSame kind of round: scenario6 min
- Your predictive maintenance system is live and the maintenance team has stopped acting on its alerts. How do you get that back?hardSame kind of round: scenario6 min
- The platform migration has no user-visible benefit. How do you rank it against features customers are asking for?hardSame kind of round: scenario5 min