The cheapest endpoints in the LLM market - sub-0.20 per million input tokens for workable model classes, and promo prices that dip below that - have persisted for months rather than weeks. A naive read calls it a subsidy. A better read is a game.
The question that muddies the analysis is that "below cost" has two different costs. Full cost includes amortized silicon, data-center build, R&D allocation - a number no provider publishes and no one can verify. Marginal cost, the cost of serving one more token on a fleet that is already running, is mostly power and a little memory-transfer overhead - and that number is verifiable. A price below full cost and above marginal cost is not a burn; it is a positive-contribution sale. A price below marginal cost is a burn, and the market tells you which case you are looking at, because burns are short.
The floor that makes it a game
You can bound the marginal cost of inference from the supply side, and the bounds are surprisingly tight:
- Power. A data center serving inference at stated efficiency puts the electricity content of a million tokens at a few dollars at scale - which is the entire vertical of the floor. Everything else is either already paid (the GPU) or almost nothing (the memory bandwidth, which the chip pays for out of sunk silicon).
- Amortized silicon. The HBM and the accelerator are bought for 3-5 years of revenue. The 2026 capacity deals (a 750-megawatt lease priced at roughly 0.47 per megatoken of reserved throughput) show the full-stack number whispered in signed leases: full cost is an order of magnitude above the electricity floor, and it is falling per token as utilization climbs.
- The persistence evidence. The fast tier at 0.20 / 1.20 has survived multiple quarters of other tiers repricing. A true burn below marginal cost gets cut in weeks - the CAC math does not survive its own bookkeeping. Survivors are, by construction, above marginal cost.
That structure - a hard floor that is easy to verify, and a ceiling of demand that is loud about itself - is exactly the shape that produces loss-leader bidding.
The bidding logic, one paragraph
Sell just above the floor when your objective is not this product's revenue but its routing share. Every request you win at 0.20 is a request the buyer's router has now classified into your price cell - and routers are sticky, because revalidating a routing rule costs an engineer an afternoon. The subsidy is not in the token; it is in the routing table. This is the same logic the airlines have run for fifty years on legacy fares: sell a seat below full cost to own the booking record, and collect the commerce in the class you actually want to sell. The exit condition is also visible: the leader raises the floor-model price once the routing share stops being worth the burn - and the whole market ratchets up in lockstep, the way airfares do after each consolidation wave.
What a buyer should take from this
Three operational consequences, all of them about not being the subsidy:
- Below-list, above-floor prices are safe. Below-floor prices are not. If an endpoint is tracking at something like a third or less of the market floor for a model class you can independently sanity-check, assume it is subsidizing - and subsidized rates come with TTLs. Build your budget for the repricing, not the rate.
- Your volume makes you visible. The providers are bidding for routing share, not customer satisfaction. If you are in their top-decile spend, you are exactly the account a repricing targets. Keep a second qualified provider warm; a warm failover is the only hedge a subsidized rate offers you.
- Marginal cost is what the industry can actually defend. Power, HBM, amortized silicon: the floor is the physics, and the physics are public. When the capacity deals put full-stack numbers in a signed lease, the loss-leader game gets shorter - a subsidized market with a visible floor is a countdown, per our understanding of per-token pricing, and the price movements tracker is where to watch the ratchet turn.
What to do
- Estimate the floor for the model classes you buy: power content of a token at stated utilization, plus amortized silicon at a 4-year life. Your number will be coarse; the market's is not.
- Price out which of your current rates sit below your estimated floor, and put those endpoints on a repricing watch.
- Keep a second provider at <100 ms of integration distance for anything subsidized the primary is serving.
- Treat a below-floor rate as routing-share math in the provider's quarter, and size your budget for the lockstep ratchet that follows.