Back to blog

Service credits are the only honest probability in an SLA: reading the fine print

The 99.9 percent is the marketing. The service credit - the percentage of fees the provider will actually pay if the service fails - is the actuarial number: the provider's own price for the failure, calibrated to the provider's own downtime data. Read the fine print and the SLA stops being a promise and starts being a probability. Here is the read.

Mappace Team · Research2026-07-017 min read
ReliabilityResearch

The SLA's "99.9 percent" is the least informative number in the document. It is a target, it is measured per the provider's own window and its own definitions, and it is the number the buyer is meant to read. The most informative number in the document is the one the buyer is meant to skip: the service credit, the percentage of monthly fees the provider will actually pay if the service fails, and the cap on that payment. The credit is the actuarial number, because it is the provider's own price for the failure, set by the provider, and the provider sets it against the provider's own downtime data. The 99.9 percent is the promise. The credit is the odds.

The actuarial read

An insurer prices a premium against the expected claim. The provider prices the credit the same way: the credit rate is the expected payout per unit of fees, and the expected payout is the probability of a claimable failure times the payout on failure. The provider sets the credit so the expected payout is a small, known fraction of revenue - because a credit that is priced wrong is a loss line, and the provider will not run a loss line on the reliability it can measure.

The read, per number:

  1. The uptime target is the complement of the allowable downtime. 99.9 percent over a month is 43.8 minutes of allowable downtime; 99.95 is 21.9; 99.99 is 4.4. The target is the downtime budget the provider has priced, and the budget is the provider's own read of its failure rate, rounded to a number that survives a sales call.
  2. The credit rate is the price of the failure, per claim. A 10 percent credit for a claimable month is the provider's maximum payment for that month's failure, as a fraction of the month's fees. The rate is set against the expected frequency of the claimable event, and a provider that runs a 10 percent credit is telling you, in actuarial language, that the claimable month is a rare event in its books.
  3. The claim cap is the price of the total failure. The cap on the credit (a percentage of monthly fees, a dollar floor or ceiling) is the provider's maximum payment for the worst case, and it is the number that tells you the contract's price of a full month of outage. Divide the cap by your actual cost of a full month of outage, and the ratio is the amount by which the contract underprices the risk relative to you - the risk premium you are carrying, for free, in the gap.

The three numbers, together, are the provider's own probability model, itemized: the frequency (the target), the per-event price (the credit rate), and the worst-case price (the cap). The SLA, read as an actuarial document, is the only part of the contract where the provider has put a number on its own reliability, and the number is the credit, not the target.

Why the target is not the probability

The 99.9 percent is not the probability of uptime, for three reasons, all structural:

  • It is a target, not a measurement of you. The provider measures the target against its own window, its own health checks, its own definition of "down." A degradation that your traffic experiences and the provider's health check does not is not in the 99.9, and the 99.9 is the provider's 99.9, not yours.
  • It is the complement of a budget, not a description of a distribution. 43.8 minutes of allowable downtime per month is a budget, and the budget is set at the level the provider can hit at the current fleet, not at the level the failure distribution actually sits. The distribution has a tail - the outage-cluster shape, where the failures come in days, not minutes - and the budget is the mean, not the tail.
  • It does not price the tail. The target is the frequency of the usual failure; the credit is the price of the claimable failure; and the tail event (the 19-day suspension, the multi-day cluster) is the one the target does not describe and the credit does price - at the cap, because the cap is the worst-case line.

The tail is the whole reliability question for a gateway buyer, because the tail is the event your failover has to cover, and the target is the number that does not describe the tail. The credit, at the cap, is the number that does - and it is the number to read.

The buyer's actuary

The read, as an operating practice:

  • Price your risk against the cap, not the target. Your cost of a full month of outage on a provider, at your volume, is a number. The credit cap is the provider's payment for that month. The gap is the risk you are carrying, and the gap is the premium you are paying, in expected loss, for the reliability the contract does not price. Size the failover (the second provider, the warm path) against the gap, not against the target.
  • Read the credit rate as a frequency signal. A provider that raises its credit rate is telling you, in actuarial language, that the claimable month has become more frequent in its books. The rate is the provider's own frequency read, and it is the one reliability signal that is calibrated to the provider's data, not to the provider's marketing.
  • The window and the definitions are the fine print that prices the target. The measurement window, the health-check definition, the "down" threshold: these are the variables that set how much of your experienced failure is inside the 99.9, and they are the variables the credit is paid against. The read is per-provider, per-term, and it is the line item the SLA summary skips.

What to do

  • For each provider in your routing, compute the gap: your cost of a full month of outage at your volume, minus the credit cap. The gap is your carried risk, and the failover budget is sized against it, not against the 99.9.
  • Track the credit rate per provider across the contract renewals. A rate change is the provider's frequency read moving, and it is the reliability signal that leads the status page.
  • Read the measurement window and the "down" definition per provider before you size the failover, because the window and the definition set how much of your experienced failure is inside the target, and the inside is the only part the credit pays.
  • Treat the 19-day suspension class (the off-switch risk event) as the tail the target does not describe and the credit prices at the cap - and size the multi-provider hedge against that tail, because the tail is the event the contract is quiet about and the cap is the only number it puts on it.