Most of 2026, the LLM market did the same thing it has been doing for three years: every tier got cheaper. The exception is the top. GPT-5.5 stepped from 2.50 / 15.00 to 5.00 / 30.00 per million tokens when it shipped in April. Claude Fable 5 launched in June at 10.00 / 50.00. Opus 5 landed in July at 5.00 / 25.00. Meanwhile the mid tier held - Sonnet 5 stayed at 2.00 / 10.00 after its announced September hike was withdrawn - and the fast-tier fell to 0.20 / 1.20 (Luna) and 0.75 / 3.75 at promo (Gemini 3.7 Flash). The frontier is the only segment where prices went up. Economists have a name for a market that does this, and it used to belong to cars.
The bifurcation pattern
Premium markets bifurcate when the top stops competing on performance and starts competing on capability:
- The base of the market commoditizes. Every year a cheaper, faster, "good enough" model replaces last year's mid tier. This is the floor moving down - it is the story of the last three years, and it is the story of 2026's mid tier.
- The top pricing detaches. When the leader's capability is the first in its class - a task the next-best model cannot do at any price - the leader is no longer selling performance per dollar. It is selling the only thing that can do the job, and that is a luxury good. The buyer compares "does it work" to "does not exist," not "does it work per dollar" to "does it work per dollar elsewhere."
The 2026 mid tier also got an architectural tailwind: the MoE families behind the GPT-5.6 line (Sol, Terra, Luna) and their rivals spread the compute cost across expert sub-models, which is why the middle can hold a 2.00 price point while the top steps up. The mid tier is the efficient frontier; the top is the capability frontier, and those two frontiers now price independently.
What the inversion does to your cost model
If the frontier is a luxury SKU, the question you should ask about it stops being "is it cheaper?" and becomes two other questions:
- Which of our workloads does only the frontier pass, and at what volume? This is the real frontier market - and it is suspiciously small in most products. If you cannot name the tasks, you are paying the luxury tier for the commodity tier's work.
- How much does the frontier save me? A 10.00 / 50.00 model is defensible if a task that 2.00 models fail in six attempts passes on the first. The correct comparison is cost-per-success across tiers - attempts_x_rate, not rate - and it should be measured per task class, not eyeballed.
Pricing governance flips to match: the mid and fast tiers should be flood-able (route everything that passes the quality bar), while the frontier tier should be rationed (explicit allow-list of tasks, spend cap, monthly re-qualification of the allow-list). That is the opposite of the old instinct, which was to ration the frontier's relatives because the frontier itself was the cheap, safe choice.
The watch items
- The promo structure of the new frontier launches (Sol at 4.00 / 20.00 through November 21) may be a test of how much the top tier can carry - a post-promo step up would confirm the bifurcation; a hold would suggest the market got deeper than the providers expected.
- A second $5+ frontier entry from a non-OpenAI/Anthropic lab would start eroding the luxury position - capability that is "first in class" is not "only in class."
- Your own frontier spend as a share of total LLM spend is the single dashboard number worth watching quarterly. It should be small and declining as a share; if it is growing, either your workload got harder (fine) or your routing got lazy (not fine).
What to do
- Run a cost-per-success benchmark for your top 5 task classes across your frontier tier and your mid tier. Measure attempts, not rates.
- Put the frontier tier behind an explicit task allow-list with a monthly budget and a quarterly re-qualification.
- Let everything else flood the midtier at rate - the floor is the part of this market that is still deflating.
- Revisit the split every time a frontier promo expires; the price movements tracker is where to watch for the step that confirms the bifurcation. For current rates across all three tiers, the pricing comparison has the full ladder.