Every cost story in inference eventually hits a wall, and the wall is not a chip. It is a power contract. The logic: the price of a token decomposes into components, and of those components, only two never disappear - the electricity the serving node draws, and the amortized cost of the rack that draws it. Everything else - software, team, support, margin - is a variable. The two permanent components are both priced on long-horizon contracts, both are public enough to bound, and their sum is the floor.
The math, with assumptions labeled
Start from decode, where the tokens come off the line, and keep every number order-of-magnitude:
- A serving accelerator at decode draws on the order of 300-400 watts under load; a node's overhead (memory, networking, cooling share) takes a meaningful fraction of that.
- At a conservative 100-150 tokens per second per accelerator on a 70B-class model, one million output tokens takes roughly 2 card-hours - give or take.
- Two card-hours at that draw is under a kilowatt-hour. At a data-center power price in the 6-10 cent range, the pure electricity content of one million output tokens is a few cents to a dime.
Two observations follow, and they are the actual point:
- Electricity is a small but unkillable share. It is not the whole cost - the rack amortization dwarfs it - but it is the one component that exists at the 1st token and the trillionth, that cannot be negotiated away by a promo, and that is priced on a 10-20 year instrument. The providers in the capacity-floor post are, in part, locking that instrument before they sell a single megatoken.
- The floor is (amortized rack + power). The rack term is the HBM, the accelerator, the site - the 750-megawatt lease's capital stack, amortized over the term. That term falls with utilization and with silicon cost, which is why the floor is a descending staircase, not a line. But it never reaches the power line: the power line is the asymptote.
Why the power contract, not the chip, sets it
The chip is a spot-good (a purchase, one direction, devalue, replace). The power contract is a forward-good (a 10-20 year PPA, or a wheel, priced before the silicon arrives, binding after it is replaced). A provider can re-forecast its silicon at every refresh cycle; it cannot re-forecast its power at every refresh. The provider that signs power cheap and locks it long is the provider whose floor is lowest - and that is a durable advantage that no model launch touches, because the model is swappable and the PPA is not.
This is also why the capacity deals are the industry's price discovery: they are the only place where the full stack - power, silicon, site, term - is written into one signed number, and the 0.47-per-megatoken-cleared is the reserve price of that stack. Spot prices sit below the reserve because they do not carry the reserve's amortization, and the distance between the two is the whole loss-leader game in the below-cost analysis.
The buyer-side consequence
You do not pick a power contract; you pick the provider that did. The practical versions of this are:
- The repricing floor is visible. When a provider's rate compresses toward a few-cents-per-electricity-content, it is not a typo - it is the provider's utilization approaching the point where the marginal cost is mostly power. That is also the point where further compression becomes a race to the power floor, and the race is the HBM and power arithmetic running below the rate card.
- The durable providers are the ones with cheap power. A provider whose power is the 4-cent category has a lower floor than the 10-cent category, at the same silicon, and the gap is the gap in that provider's repricing resilience. It is the one durable competitive advantage in the stack that is not swappable.
- Your marginal cost proxy is electricity, not minutes. When you model "what if this gets 10x volume," the cost term that scales is the power term, and the term that does not is the amortization - the opposite of what the spec sheet teaches.
What to do
- Compute the electricity content of your 1M output tokens on your actual model class and accelerator assumption, and put that number beside your current rate - it is your proxy for the provider's marginal cost, and the delta is the provider's contribution margin.
- Watch the capacity contracts as they land. Each one is a standardized unit of the full-stack floor, and the industry's long-term price is set where the reserves clear.
- When a provider reprices below your margin expectations, check the repricing against the power content: a reprice that is still 10x the electricity content is a margin story, not a floor story - and it does not need to be your problem.