Skip to content
Preptima
hardDesignCase StudyMidSeniorStaffLead

Underwriting want to put a new external data source into the price. What has to be true before it goes anywhere near the rating algorithm?

It has to be available at quote time for a known share of risks, obtainable as it stood historically so a backtest means anything, defensible as a factor rather than a proxy, and stable enough that the vendor cannot change your filed rate for you. Everything else is a modelling detail.

6 min readUpdated 2026-07-29Target archetype: Enterprise Captive, Product Startup
Practice answering out loud

What the interviewer is scoring

  • Whether point-in-time availability of the historical values is demanded before any backtest result is believed
  • Does the candidate ask what the match rate is, and what the price is for a risk the vendor cannot match
  • That the vendor's freedom to change its own methodology is identified as an uncontrolled change to a filed rate
  • Whether lift is tested over the factors already in the model rather than measured in isolation
  • Does the answer cover exit, and what happens to a filed rate when the data contract ends

Answer

The modelling question is the last one to ask

The conversation almost always opens with predictive power, and predictive power is the cheapest thing to establish and the least likely to be the reason the project fails. A data source that discriminates beautifully between good and bad risks is useless if it cannot be fetched inside your quote window, if it only exists for two thirds of your applicants, if the version you tested on is not the version that will be available at the point of sale, or if you cannot explain to an examiner what it measures. Each of those is a hard constraint, and each of them kills more proposals than a weak Gini coefficient does.

So the order of the assessment matters. Establish that the data can be obtained, at the right moment, in a form you can stand behind, and only then ask what it is worth.

Can you get it when you need it, and for whom

A rating factor has to be present at quote. That sounds trivial until you put it next to the response-time budget of a distribution channel where a slow answer is no answer at all. Every external call in the quote path has a latency distribution with a tail, an availability figure below one, and a price per call, and adding a factor means accepting all three.

The harder half is coverage. Vendors match on identifiers, and matching fails: an address that does not resolve, a vehicle not in the reference file, a person with no history at that address. Ask for the match rate against your own book rather than against the vendor's universe, because they are different populations, and ask how it varies by segment. Then confront what you will charge a risk the vendor cannot match, because that default is itself a rate. Set it optimistically and every unmatched risk becomes cheap at your expense. Set it pessimistically and you are surcharging people for the vendor's data-quality problem, which is both unfair and, in the aggregator channel, immediately visible as a competitive hole. And if the default is visibly different from the matched price, you have created an incentive to present the risk in a way that fails to match, which is an arbitrage somebody will find.

Historical values, as they stood then

This is the technical trap in the whole exercise and the one that quietly invalidates most first attempts at a business case. To know what the factor is worth you have to attach it to policies you have already written and observe the losses that followed. That requires the value the vendor would have returned at the time the policy was quoted, not the value they hold today.

Where the data is a snapshot of something that changes, using today's value is not a small approximation. A credit-derived signal, a claims-history count, a business turnover figure, a property attribute updated after a renovation: each of these has partly been shaped by the very outcome you are trying to predict. A claim you paid may have caused the value you are now using to predict it. The result is a backtest that looks extraordinary and a live model that does nothing, and by the time you discover that, the factor has been filed and the rate change has been taken.

The remedy is to insist on a point-in-time extract, dated to the quote date, and to be explicit in writing about what you can and cannot conclude if the vendor cannot provide one. Sometimes they cannot, and the honest fallback is a forward test: take the factor live in shadow mode, store the value returned at every quote alongside the quote, and wait for the experience. That costs time and buys credibility, and it is the same discipline as pinning a quote to its rate version, applied to a factor instead of a table.

Is it new information, or a restatement of what you already price?

Lift measured on its own is not the question the pricing exercise needs answered. Almost any plausible external signal will separate good risks from bad, because so much is correlated with so much else, and much of it is already in your rate through postcode, age, vehicle group, tenure or sum insured. What matters is the incremental lift once those are in the model, on a holdout drawn from the risks you actually write rather than from the market as a whole.

That distinction is worth pressing because your book is already selected. The factor's power on a broad sample may come almost entirely from a segment your underwriting rules decline anyway, in which case you have bought a variable that will do nothing to your loss ratio while adding a vendor, a cost per quote and a latency tail. A small incremental lift on your own written business is a better result than a large gross lift on somebody else's population.

Defensibility, not just fairness in the abstract

A rating factor has to survive being described out loud. Some jurisdictions prohibit specific characteristics in insurance pricing outright, and where they do, a proxy for the prohibited characteristic is generally no more acceptable than the characteristic itself. External data is more exposed here than internal data because it is often an opaque composite: a score whose construction the vendor treats as proprietary can be carrying signal you would not be allowed to use directly, and you cannot rule that out by inspection.

Two practical consequences follow. You need to know enough about the construction of the input to state what it measures in plain language, which is a contractual requirement on the vendor rather than an analytical one on you. And you need to test the fitted effect for concentration against characteristics you are not permitted to price on, so that the answer to a regulator's question is evidence rather than assurance. In a filed market this all becomes concrete anyway: the factor and its relativities go into the filed algorithm, the filing has to describe them, and the examiner is entitled to ask.

Whose rate is it once the vendor can change it?

The last constraint is the one most often left out, and it is the reason to treat a data purchase as an architectural commitment. Once the factor is in a filed rate, the vendor holds a lever on your pricing. If they recalibrate their score, extend their reference file, or change their matching logic, your premium distribution moves without any change on your side and without any filing, and the first evidence is a shift in conversion or in average premium that somebody has to trace back to a supplier release note.

That has to be engineered against explicitly. Version the vendor's output as part of the rate version, so a quote records which vendor release priced it. Monitor the distribution of the returned values and the match rate continuously, and alert on the shape of the distribution rather than only on errors, because a recalibration is a successful response with different numbers in it. Negotiate notice of methodology changes and the right to obtain the prior version for a period. And decide in advance what happens at the end of the contract, because a filed rate that depends on a data source you no longer buy leaves you re-filing under time pressure with a default factor doing the work of a real one.

The question is never whether an external variable predicts losses. It is whether you can obtain it at quote for a known population, reconstruct it as it stood in the past, explain what it measures, and stop somebody else's release note from repricing your book.

Likely follow-ups

  • The vendor can only supply today's value for a risk you wrote three years ago. What can you still learn from that, and what must you stop claiming?
  • The match rate is seventy per cent. How do you price the remaining thirty without opening an arbitrage?
  • How would you demonstrate to a regulator that the factor is not a proxy for something you are not permitted to use?
  • The vendor recalibrates their score and your average premium moves overnight. What should have caught that?

Related questions

underwritingratingexternal-datamodel-governancerate-filing