Two prices for the same compute
A Claude Pro plan and an Anthropic API key buy inference from the same lab, on the same GPUs, at prices that are not close to each other. Almost every AI product is built on the expensive one.
There are two ways to buy frontier inference, and the labs price them for two different buyers.
The first is a consumer plan. Claude Pro or Max, ChatGPT Plus or Pro: a flat monthly price, and it comes with the agent CLIs that run on it: Claude Code, Codex. It is priced on average usage across a population, which means the light users pay for the heavy ones, and it carries strategic weight for the lab selling it. Consumer plans are how a lab wins the habit. They are marketed, discounted, and made steadily more generous, because the thing being bought is not really the compute.
The second is an API key. Tokens, priced at the margin, sold to companies who will resell them. No subsidy, no cross-subsidy, no strategic discount. This is the price of the thing itself, and it is the one every AI product pays.
The gap is not a rounding error#
The two prices are not variations on a theme. A person on a $20 consumer plan can run agent workloads all month that would cost hundreds of dollars of API tokens to serve, and the lab is fine with that, because the population average holds even when individual users are wildly unprofitable.
You cannot get that price. It is not available to you at any volume, and it is not going to be, because the whole point of the subsidy is that it is attached to a person rather than to a business. What you can do, and the thing this entire company exists to do, is stop buying the expensive one on your users' behalf when a large fraction of them are already holding the cheap one.
The subsidy is attached to a person. It is not for sale. But the person can bring it with them.
Follow one $20 through#
Here is the same twenty dollars, put down on either side of the counter.
Spent on a Claude plan, all of it buys compute at a price set for a person and subsidised by the lab selling it.
Paid to you, it is about $19 after payment fees. What is left has to buy inference at the full API price, and also fund your servers, your salaries, your support, and whatever margin makes the business worth running. The user ends up with a fraction as much of the identical product, from the identical provider.
This is not a pricing mistake you can optimise your way out of. It is structural. Every AI app that resells credits is competing against its own supplier's consumer tier, using its supplier's most expensive price, and paying Stripe on the way through.
Why it keeps getting worse#
The gap is not stable. It widens.
Consumer tiers are a competitive front: every time one lab makes its plan more generous, the others follow within a quarter, because losing the habit is worse than losing the margin. API prices move too, but they move slowly and they move for everybody, including your competitors. So the ratio between what a person can buy for $20 and what you can buy for $20 has been drifting in one direction for two years, and there is no mechanism that reverses it.
If your unit economics depend on that ratio staying where it is, they depend on the one number in this market that is reliably getting worse for you.
What you can actually do about it#
Three options, honestly stated.
Eat it. Buy tokens, resell credits, price high enough to survive your heavy users. This works, and it is what most products do. It caps how generous your free tier can be, it makes every power user a liability, and it means your gross margin is set by a supplier who also sells directly to your customers.
Ration it. Message limits, credit packs, "fair use" ceilings. Same economics, moved into the product surface, where the user can feel it. You have converted a margin problem into a retention problem.
Route to the compute the user already owns. A user who has a Claude or ChatGPT plan is holding subsidised capacity that is idle most of the day. If your app can run on their plan, on their machine, those runs cost you nothing to serve, and you can hand some of that saving back as a bigger allowance, which is a better offer than any competitor paying API rates can match.
The third one is not free. Someone has to route the run, decide when it is allowed, meter what it cost, and bill the user at your price rather than at cost, and do all of it without your call site knowing which side of the line it landed on. That is the layer we build. The quickstart is one changed base URL, and Who pays is the part where the accounting stops being hand-wavy.
The argument for doing it is just this: you are currently buying, at the highest price anyone pays, a thing that a large share of your users already own.