Back to blog

Model EOL is an unpriced risk: the actuarial premium you already carry

Every model you have routed has an unannounced end date, and the migration it forces is a cost you already pay, in a line item nobody named. Model deprecation is an actuarial risk with a survival curve, a premium, and a hedge - and the premium is the re-validation cycle your team runs on a schedule no one wrote down. Here is the math that makes it visible.

Mappace Team · Research2026-07-287 min read
ReliabilityCosts

The model you are routing at has an end date. It is not on the rate card, it is not on the changelog's front page, and it is not in the SLA's failure definitions, because a deprecation is not a failure - it is a planned retirement, with a notice period (a 30-day class is the standard, and the providers adjust the notice as the cadence moves). But the retirement forces a migration, and the migration has a cost, and the cost is recurring, on the model's schedule, in a line item that most cost models do not name.

The actuarial frame is the correct one, because the model's retirement is the same risk shape the insurance industry prices: an asset with a survival curve, a replacement cost, and a premium that is the expected replacement cost weighted by the retirement's frequency. The premium is not a line item. It is the re-validation cycle your team runs, and the cycle is the premium, paid whether or not anyone has priced it.

The survival curve, and who holds it

Every routed model has a survival curve, and the curve is the probability that the model is still in production, at full support, as a function of time since launch. The curve is not public in the actuarial sense (there is no published mortality table for model generations), but it is estimable, from the deprecation history that is public:

  • The launch-to-deprecation interval, per model family, across the providers, is the observation set. The interval has a shape (a median, a tail), and the shape is the survival curve's estimate, per family, per provider.
  • The notice period is the curve's hazard, made operational: the 30-day class notice is the window between "the model is being EOL'd" and "the model is gone," and the window is the migration's budget, in calendar time.
  • The cadence is the curve's frequency: the number of EOL events, per year, across the models you route at, and the cadence is the frequency term in the premium, the same as the claim frequency is the frequency term in the insurance premium.

The three - the interval (the curve), the notice (the hazard window), and the cadence (the frequency) - are the actuarial inputs, and all three are estimable from the public deprecation record, per provider.

The premium, itemized

The migration's cost, per EOL event, is the premium per claim, and the itemization is the part the cost model misses:

  1. The re-validation. The replacement model has to be validated against the task's quality bar, on the task's eval set, before it is routed at. The re-validation is the eval run, the side-by-side, the pass/fail determination, and the cost is the engineering time of the run, plus the eval infrastructure, plus the decision. It is a fixed cost per task class, per EOL event.
  2. The prompt and the tuning. The replacement model's prompt sensitivity is different from the incumbent's, and the prompt that passed at the incumbent can fail, or degrade, at the replacement. The tuning is the prompt rework, the parameter re-tuning, the temperature and the top-p and the system-prompt rewrite, and the cost is the same order as the re-validation, per task class.
  3. The regression and the canary. The routed-at-replacement model has to be canaried (a traffic slice, a comparison against the incumbent's output on the shared tasks) before the full cutover, and the canary is the risk control, with its own cost - the slice's monitoring, the comparison, the rollback path held warm for the canary window.
  4. The dual-run carry. The incumbent and the replacement run in parallel for the validation and the canary window, and the dual-run is the carry cost - two providers' invoices, for the window, at the full volume. The carry is the one line item that scales with your volume, and it is the line that grows with the traffic the EOL happens on.

The premium per event is the four summed. The premium per year is the premium per event times the expected EOL events per year across your routed set - the frequency term from the survival curve's cadence. The number is not on any invoice, and it is the number that the routing decision implicitly pays, because a routing that concentrates at one model family is a routing that concentrates the premium, and a routing that spreads is a routing that diversifies the claim.

The hedge, and why the abstraction is the policy

The actuarial hedge for a replacement-cost risk is the one that reduces the per-claim cost, and for a model EOL the per-claim cost is the re-validation plus the tuning plus the canary, and all three scale with how different the replacement is from the incumbent, and how unmeasured the task's quality bar is. The three hedges, in order of leverage:

  • The eval harness is the policy. A task class with a maintained eval set (the golden tasks, the pass bar, the regression suite) has a re-validation cost that is an eval run, not a project. The harness is the thing that converts the re-validation from a one-off engineering effort into a scheduled measurement, and it is the single highest-leverage hedge, because it reduces the per-claim cost at every claim, not the frequency.
  • The abstraction layer is the diversification. A routing that sits on an abstracted model interface (the gateway's one-endpoint shape) has a cutover that is a config change, not a code change, and the cutover's cost drops by the code-change term. The abstraction does not reduce the re-validation (the eval still runs), but it removes the integration cost from the claim, and the integration cost is the part that scales with the team's other work.
  • The multi-model design is the frequency hedge. A routing that qualifies two models, per task class, at the quality bar (the routing strategies second-provider design) is a routing where the EOL of the incumbent is a re-weight, not a migration, because the replacement is already qualified, already canaried, already invoiced. The dual-qualification is a standing carry cost (the second model's warm-path spend), and the carry is priced against the avoided claim, the same as the failover carry is priced against the avoided outage.

What to do

  • Estimate your survival curve, per routed model family, from the public deprecation record: the launch-to-EOL interval, the notice period, and the cadence, per provider. The three numbers are the actuarial inputs, and they are the ones the cost model is missing.
  • Itemize the premium per event (the re-validation, the tuning, the canary, the dual-run carry), per task class, and put the annual premium - premium per event times the expected events - into the LLM cost model, as a named line, next to the token spend.
  • Build the eval harness for the task classes that carry the highest EOL exposure (the classes where the re-validation is a project today), because the harness is the hedge that reduces the per-claim cost at every future claim.
  • Qualify a second model, per high-exposure task class, at the quality bar, so the EOL is a re-weight and not a migration - and price the warm-path carry against the avoided claim, the same actuary's math the service-credit read applies to the reliability risk.