An app declares which cost classes it will accept, not which runtime to use, which is what lets one API back a free-forever app and an enterprise app without either changing code.
The obvious design is a runtime field: name the thing you want. It is also the
wrong one, because it forces the app to know the topology: whether this user has a
machine paired, whether it is awake, what is installed on it.
route lets an app declare intent instead.
JSON
{
"route": {
"prefer": ["user-hardware", "user-subscription"],
"allow": ["user-hardware", "user-subscription", "app"],
"maxCostMicros": 250000
}
}
| Field |
|
prefer |
Ordered preference. The first reachable class wins |
weights |
A mix over prefer, in relative numbers. Omitted means strict order |
allow |
Classes usable at all. A hard ceiling. Omitted means anything Soba can reach |
maxCostMicros |
Ceiling on what one run may cost. Metered classes only. On app it is passed to your fallback, the only place it can be enforced |
Most apps never write route by hand: a plan compiles to one on every
run, so pricing and routing stay the same decision.
One API, several business models#
free-forever app prefer/allow ["user-hardware", "user-subscription"]
bring-your-own + "user-api-key" they pay their provider; you pay nothing
prosumer app + "open-weights" a sleeping laptop falls back instead of failing
enterprise app + "frontier"
None of these change any code. They change two arrays.
user-api-key is worth a second look on that list, because it is the only class that
costs the user money and costs you none. An app that allows it serves people who
have a provider key but no subscription and no local model, and it serves them without
touching your bill. An app that means "free to everyone who uses it" should leave it
out, which is exactly why it is not folded into user-hardware.
allow is enforced twice, and that is not redundant#
It is enforced when a machine is chosen, and again on the machine when the run arrives.
A machine can serve a class the app did not ask for: an owner who sets
allowMeteredKeys turns their CLI into user-api-key. So an app that declared which
wallets it would accept must not have a run land on another one because it was routed
wrong. This is the same reasoning that puts
the policy intersection on the machine.
What happens when a step fails#
A step that is not there is skipped before the run starts: an asleep machine, a
missing credential, a class the route bars. The run goes to the next step the
order allows.
A step that breaks after the run has started is newer, and narrower. If a
machine drops, crashes or is cancelled before the run has produced any output,
the run falls to bought supply once, provided the route allows buying and the
app is inside its allowance and period cap. If none of those hold, the run ends
with the machine's error.
Once output has reached the caller, nothing is retried. A second answer appended
to half of an abandoned one is worse than an error, because the caller cannot
tell it happened.
Two paths are deliberately excluded. A connected subscription that refuses ends
the run rather than buying the answer instead — see
the ChatGPT channel. And bought supply has its own
retry: runMetered walks its offer list across suppliers and both bought
classes, stopping the moment one has streamed.
Splitting traffic#
prefer is an order. weights turns it into a mix.
JSON
{
"prefer": ["user-hardware", "open-weights"],
"weights": { "user-hardware": 70, "open-weights": 30 }
}
The numbers are relative, not percentages of your traffic, and this is the
part worth reading twice. They are normalised against the classes that user
can reach at the moment the run arrives:
user with a machine connected 70/30 between their machine and bought supply
user with nothing connected every run bought — one candidate, one winner
So a 70/30 split does not mean 30% of all runs are bought. It means 30% of the
runs by people who could have gone either way. Users with one option send
everything there, whatever the split says.
Three more properties:
- Drawn per run. Each run is an independent draw, not a rotation. Ten runs on
a 70/30 can land eight and two; the ratio is a limit, not a quota.
- The draw picks the first attempt only. Whatever wins is what the run tries.
- A failure does not re-draw. See below.
What happens when a step fails#
Reachable is not the same as working, and the ladder only decides the first.
- A machine that drops mid-answer ends the run. Nothing is re-routed onto
compute you pay for on your behalf.
- A connected subscription that refuses ends the run, for the same reason: an
app that priced itself on its users' own compute should hear about it rather
than quietly be billed instead.
- Bought compute is the exception. It walks your suppliers in order until one
starts answering, and stops there — half an answer and an error beats half an
answer followed by a different whole one.
Which machine, when a user has several#
prefer chooses the class. Within user-hardware, Soba picks the user's least
busy reachable machine, where reachable means online or busy with capacity to
spare. Losing a machine sets its capacity to zero, so an offline one drops out
without being named.
Naming a runtime anyway#
You can still pin a run to claude-code or codex when you have a reason to. Omitted,
the machine uses its default. Prefer route unless you specifically need the runtime.
The cost ceiling is checked between turns#
maxCostMicros stops a run that reaches it. The check happens between turns because
nobody knows a turn's cost before it happens: the guarantee is "stops as soon as it
knows", not "never exceeds".