Soba is the LLM layer for agentic applications. You replace your LLM calls with Soba, and it handles the pricing plans, the billing, and connecting each user to compute they already pay for.
It takes over everything those calls drag behind them: what a run is allowed to use,
what it costs, who pays for it, and what you charge for it.
One plan, as a user sees it.
The unit of all of that is a pricing plan: the one your users see, with a price on it.
You write it once: what it costs, what it includes, what happens past the allowance, a
ceiling on any single run, a trial. In the same object you say which compute those runs
may use, because pricing and routing are one decision. Soba compiles it into real
prices in your own Stripe account, runs the checkout, and bills each user at your rate.
See Plans and billing and Stripe.
One layer, slid in underneath. Above it nothing about your app changes; below it, a run lands on whatever compute that user’s plan allows — usually one they are already paying for.
What makes that rate yours to set is where the compute comes from. A user who connects
their own Claude or ChatGPT plan runs on it, on their own machine, and those runs cost
you nothing to serve, so the plan can pay them back for connecting, with a larger
allowance or with credit. Everyone else is handed back to your app, which serves them
on your own keys exactly as it does today, on the same plan. Either way it is one bill
at your price, collected in your own Stripe account, and what the run cost to serve
stays your margin.
None of it is yours to operate. Every run is logged with what it cost and who paid for
it, in micros; usage per user, connected machines, spend ceilings and failures all sit
in one dashboard, whose lead number is how much inference you did not pay for this
month.
Why this exists#
Every free user of an agentic app costs you money, because every run is inference you
pay for. So the free tier gets rationed: a few credits, the good model behind the
paywall, a cap that lands on day three. Most users hit that wall before they've worked
out what the product is for, and someone who hasn't built the habit has no reason to
pay.
A longer, more generous trial is the obvious fix, and the one most teams can't afford.
Its cost grows with how generous it is and how many people take it.
But many of your users already have the compute to run it. Anyone on Claude Pro or Max,
or ChatGPT Plus or Pro, has a monthly plan with the agent CLIs included, installed on
their own machine under their own login. Soba lets your app run on that during the free
period. It handles the connect step, the routing, the policy that keeps each run within
the terms that plan came with, and the meter that tells you what each user used.
That means the trial can last until the product is part of someone's week, without the
inference bill landing on you. You set the envelope: how long, how much, which
features. When it ends, the user moves onto your paid plan, and that plan should offer
something their own subscription can't. Soba is the bridge to that moment, not a
replacement for charging.
The same trial, twice. Paid for in credits it ends before the habit forms; run on compute the user already has, it lasts until the habit is worth paying to keep.
The run has to happen on their machine#
"Unless previously approved, Anthropic does not allow third party developers to
offer claude.ai login or rate limits for their products, including agents built on
the Claude Agent SDK."
A run on the user's own machine, under their own login, is a person using the plan they
bought. Moving that same run into your cloud is what the wording above addresses. That
is why the run lands where it does, and why attribution is
enforced in code rather than promised.
Whether any particular deployment is permitted is between you, your users and the
provider, and the providers are the only ones who can say. Soba is built to keep the
distinction above true, and Provider terms says exactly what it
does to hold it.
The five surfaces#
One layer, five surfaces. The routing everyone notices first is only one of them.
One endpoint, two destinations: a user’s own AI once they have connected it, your own app until then. Decided per run, never at your call site.
| Layer |
What it covers |
| Route |
Which compute serves the run. An app declares the cost classes it will accept, not a runtime, so one API backs a free-forever app and an enterprise app without either changing code |
| Connect |
How the user's access gets there: a Claude or ChatGPT plan, a local model, or a key. An OAuth popup or a paired machine; no credential is ever pasted |
| Enforce |
What policy binds it. A run gets the overlap between what the machine's owner allows and what your app asked for, decided on the machine so a compromised control plane cannot escalate |
| Meter |
What it costs. Per-user usage stamped with its cost class, in integer micros, with a spend ceiling checked between turns |
| Bill |
What your users pay you. Plans you build in Soba become prices in your own Stripe account, and checkout runs through Soba so every payment is linked to the trial that led to it. Charged at your price, which may be nothing |
Behind them sits a dashboard whose lead metric is deliberately not a usage graph:
inference you didn't pay for this month, auditable against your own provider
bill, and a number no gateway can compute.
Adopting it#
Nothing beyond the first is required.
- Swap the URL. One line. Your SDK, plans, billing and auth all
stay. You get metering, ceilings, and a clean handback to your own code for anyone
with nothing connected.
- Add a component.
<ConnectCompute /> in your settings page.
Users connect; a local model or a key of their own starts serving runs with no change
at the call site. Their Claude or ChatGPT plan is reached through
soba.run(), which owns the agent loop chat/completions cannot.
- Adopt the plan layer. Pricing, checkout on your own Stripe
account, entitlement, dashboard.
The first asks for a base URL, not a rewrite, and everything after it is optional.
Most of the value arrives there, before any of your users has connected anything.
What it is not#
Not an inference reseller. Soba owns no GPUs, sells no tokens and never holds your
provider keys. A run lands on compute your user already pays for, or is handed back to
your app to serve the way it does today.
Not a payment processor. Billing runs in your own Stripe account, and Soba never
holds or moves your money. See Stripe.
Not agent identity. Soba builds the authority layer because compliance requires
it; the identity category is a standards fight won by consortia.
Not a consumer brand. The end user's relationship is with your product, not
with Soba.
Not a subscription pool. A run executes on the end user's own machine, under
their own login, for their own account. Pooling one operator's subscriptions to serve
strangers is the prohibited thing; see Provider terms.
Where to start#