newSoba is in private beta. Request access
← All posts

Let your users bring their own Claude or ChatGPT plan

What it actually takes to run an agent on compute your user already pays for: the routing, the permissions, the metering, and why the call site should never know which side it landed on.

Andrea ChelloFounder, Soba
SOBA 001 ONE BASE URL

"Let users bring their own key" is an old idea and a bad one. You get a settings page, a textarea, an encrypted column, and a support queue full of people who pasted the wrong string. The user still pays API rates, you still get blamed when it breaks, and nothing about your economics changed.

What is new is different, and worth being precise about: a user with a Claude Pro or Max plan, or a ChatGPT Plus or Pro plan, is already paying for an agent that runs on their own machine. Claude Code and Codex ship with those plans. If your app's runs can execute there, they cost you nothing to serve, and the user is not paying twice: they already bought it.

This post is about what stands between that sentence and a working system.

What the call site should look like#

Start from the end, because it constrains everything else. Here is the whole integration:

JavaScript
import OpenAI from "openai";

const soba = new OpenAI({
  baseURL: "https://api.soba.so/v1", // the only line that changes
  apiKey: process.env.SOBA_KEY,
});

const response = await soba.chat.completions.create({
  messages: [{ role: "user", content: "Refactor this module" }],
  model: "auto",
  user: session.user.id,   // ← the one addition
});

That is it, and the design constraint hiding in it is the important part: nothing here says where the run executed. Not a branch, not a flag, not a second code path for connected users.

If your call site has to know, you have built a fork in your application that will be there forever, and every feature after this one gets written twice. The whole job of the layer underneath is to make "ran on the user's Claude Max plan" and "ran on a metered API tier" the same thing to the caller, differing only in what shows up on the bill.

user is a hint in the OpenAI API and load-bearing here: it is how a run is attributed to a person, and therefore how the system knows whose machine it may reach.

The four problems underneath#

Your app calls Soba at one base URL. Soba sends the run to the user's own compute if they have connected some and hands it back to your app if they have not, and calls back into your app to run tools. Your app Soba Route Enforce Meter Bill Their own compute Costs you nothing Back to your app At your price @soba-so/sdk calls your tools
One endpoint, two destinations: a user’s own AI once they have connected it, your own app until then. Decided per run, never at your call site.

Reaching a machine behind a home router#

The user's laptop has no public address, no stable IP, and a router that will not forward anything. So the connection cannot be opened from your side. A small worker on the user's machine dials out to the broker and holds one WebSocket; every frame of the run (the prompt going down, the tokens coming back, the tool calls going both ways) crosses that one socket. The end that opens it is the one behind the router, which is the only arrangement that works without asking a user to configure anything.

Then it has to survive: sleep, wake, network changes, a closed laptop lid. A worker that requires the user to restart it after every suspend is a worker that gets uninstalled in week two. See keeping it running.

Deciding what a run is allowed to do#

Two parties have opinions and neither gets to win outright. The machine's owner grants a set of permissions: what a run may read, write, execute. Your app requests a set. What a run actually gets is the intersection:

Two resolutions, each an intersection. Permissions are the machine grant intersected with the broker request; the tier is what the machine can serve intersected with what the app declared it would accept. Neither result can be wider than its inputs. The machine grant what the owner allows The broker request what your app asks for permissions never wider than either ∩ What this machine can serve the tiers it actually has What the app will accept route.allow tier one of those two lists ∩

The same shape applies to routing. The machine advertises which tiers it can serve; your plan declares which it will accept; the run goes to something in both lists or it does not go. A resolution that can only ever narrow is the property that makes this safe to reason about: no combination of a permissive machine and an eager app can produce something wider than either one allowed.

Knowing what a run actually cost#

"Free" has more than one reason, and only some of them are free to the user. A run on the user's own hardware, a run on their subscription, a run on an open-weights model you bought wholesale, and a run on a frontier API are four different economic events. Collapsing them into a boolean, which is what we tried first, makes the meter lie.

TypeScript
type CostClass = "user-hardware" | "user-subscription" | "open-weights" | "frontier";

And it has to be asked, not assumed. A Claude CLI on the user's machine might be signed in against a subscription or against an API key, and those are opposite answers to "who pays". The worker asks the CLI, cheaply and locally, rather than guessing:

Terminal
claude auth status --json

If the CLI turns out to be running on a metered key, the run reports frontier and gets metered, so the meter and the wallet agree. Guessing here is the exact failure the field exists to prevent. The details are in Who pays.

Billing at your price, not at cost#

This is the one people skip, and it is the one that makes the whole thing a business rather than a discount.

Metering is cost accounting: what the run consumed, in integer micros, stamped with which economics served it. Billing is a different number and it is yours. A user on your $20 plan is billed $20 whether their runs cost you nothing or cost you real money. What a run cost to serve is your margin, not your customer's concern, and not something they should be able to infer from their invoice.

Which is also why the pricing plan and the routing policy are one object rather than two. What a user pays and what their runs may execute on are the same decision, made once:

A pricing plan: Pro, twenty dollars a month, a thousand runs included, two cents a run after that, and unlimited runs while the user is connected to their own AI. Pro $20 / month 1,000 runs included $0.02 a run after that Unlimited while connected
One plan, as a user sees it.

The plan says the price, the allowance, what happens past it, the ceiling on any single run, and which cost classes it permits. It compiles to a route on every request. See Plans and billing.

The reward is the actual product decision#

Everything above is plumbing. This is the part that changes your funnel.

Once a connected user costs you nothing to serve, you can pay them for connecting, and the size of that payment is the lever the whole model rests on:

  • none: nothing changes, you just keep the margin.
  • included: a larger allowance while connected.
  • unlimited: the cap comes off entirely while they are connected.

unlimited is the interesting one, because it is an offer your competitors structurally cannot match. Unlimited runs while you are connected to your own AI is honest, it is enforceable, and it costs you less than the metered tier it replaces. A company buying API tokens can only reach that offer by losing money on it.

It also inverts the usual relationship with power users. The heaviest user on the platform, connected, is now your cheapest, and the one with the strongest reason to stay, because their allowance is tied to a connection they set up themselves.

What it does not solve#

Being straight about the edges:

  • Not every user will connect. Plenty have no subscription and no interest in one. They fall through to your metered tier on the same plan, at your price, and the integration does not change, which is the point of the run not knowing where it landed.
  • Their machine can be off. A run that needs a machine that is asleep has to fail cleanly or fall back, and your product needs to have decided which. That is a route decision, not an incident.
  • Provider terms are the provider's. What a consumer plan permits is written by the lab selling it, it is not the same across labs, and it changes. We track it in provider terms rather than pretending it is settled.

None of those are reasons to keep buying tokens for users who already own them. They are reasons to build the fallback properly, which you needed anyway.

The quickstart is four steps and the first one is an environment variable.

Keep reading

Stop buying tokens
your users already own.

Point your OpenAI client at one base URL. Soba prices the plan, routes the run, and bills the user at your rate.

Claude Code, Codex, Gemini CLI and Ollama are the products of their respective owners. Soba is an independent tool, not affiliated with or endorsed by any of them, and each run stays on a machine dedicated to one user, signed in with that user's own account and under that provider's terms.© 2026 Soba