The path a run takes from your app to a subprocess on someone else's laptop and back, and why the machine dials out rather than listening.
A run passes through four things, and the whole design turns on one connection between
two of them.
Every frame in the sequence below crosses that one socket, and the end that opens it is the one sitting behind a home router.
Your app is your product: your prompt, your tools, your users.
Soba decides where the run is served, meters what it used, and bills your users
for your plan through your own Stripe account.
The worker is a small program running on your user's own machine, which is the
one placement the product cannot trade away. The next section
says why.
The runtime is what actually does the thinking: their Claude Code, their Codex, or
a model endpoint.
Your app never talks to a worker, and Soba never runs a model. Each one hands off to
the next.
Why it runs on their machine#
The compute Soba is reaching for is a Claude or ChatGPT plan your user already pays
for. Plans like that are sold to a person, at a price that only works because it is
sold to a person, and they are licensed accordingly. Anthropic does not permit a third
party to offer claude.ai logins or rate limits to its own users, so there is no version
of this that runs in your cloud on your users' behalf.
What is left is the honest version: the run happens where the plan already lives. On
their machine, under their login, billed to their account. Nothing is pooled, nothing
is resold, and no credential is ever handed to you or to Soba.
Move that same run onto a server you operate and it stops being a person using the plan
they bought. That is the prohibited thing, and it is the line the worker exists to stay
on the right side of. Provider terms has the invariant and what
enforces it.
The other half of the answer is economic. A subsidised consumer plan gives your user
more compute than you could afford to buy for them, so a run served on their machine
costs you nothing and outperforms what your budget would have
stretched to. The compliance argument says it must be there; the economics say you
would want it there anyway.
Why the worker calls Soba, and not the other way round#
Soba cannot call a laptop. A machine on home wifi has no address the internet can
reach; it sits behind a router. To dial in, every one of your users would have to open
a port on that router: a support burden on you and a hole in their network.
So the worker places the call, the way a browser calls a website, and that connection
stays open. When work arrives, Soba sends it down a pipe that already exists. It is the
difference between an integration your users have to configure and one they install.
What that buys:
- no inbound port, so nothing on the machine is reachable from outside
- no public address and no port forwarding
- NAT traversal, which is what makes a laptop on a home network a viable place to
run anything at all
What happens during a run#
1. The machine says what it can do. The moment it connects, long before any run
exists, it reports every runtime it can actually serve: its models, its native tools,
its cost class. A runtime the user is signed out of is left off
that list rather than advertised as broken, so it cannot be routed to by mistake.
2. Work arrives. The prompt or the structured messages, an optional system prompt,
and two optional requests: what this run should be allowed to do, and which
cost classes you will accept. Both are requests, and the next step is why that word is
doing work.
3. The machine decides what this run may do. Nothing is spawned until two questions
are settled, and each is an intersection. Read ∩ as only what is in both:
Neither line can grant anything. The machine's owner has already said what a run may do
there, your app asks for something, and the run gets the overlap, never more than the
owner allowed, never more than your app asked for. The second line settles which
compute serves it the same way.
A run that comes out empty on either line is refused before a process starts. This is
the whole security model, and it is decided on the machine
rather than by Soba for one reason: it has to hold even if Soba is lying.
4. Output streams back. Every runtime is normalised onto
five event types. The first event names the tier that won,
before any output. A cheaper tier covering for a sleeping laptop will visibly
underperform, and a visible downgrade beats a silent one.
5. Your tools, if you defined any. The model's tool call comes back to your app,
your app executes it, and the result goes the other way. The code and the data stay on
your side throughout. See App-defined tools.
6. A question, if a tool needs consent. Only for tools inside the machine's askable
set. Anything outside that set is refused without asking anyone, because asking would
imply it could be allowed. See Approvals.
7. The run ends. Carrying its usage: the model, the provider, the token counts, and
for metered classes the cost in micros and the cost class that produced it. A usage
record is self-describing, so a meter never has to remember what it routed to in order
to know whose money was spent.
What this leaves out#
Two of those steps stand in for a page each, and neither is summarised well by one
paragraph:
- What the machine spawns at step 3. An agent runtime (
claude-code, codex) runs
its own loop, so the resolved policy is passed to it and is advisory; a model
runtime is driven by Soba's own loop, so the same policy is enforced in-process,
immediately before the tool runs. Agents and models.
- How the machine was picked at step 2. The workers that user has connected, a
preference order across tiers, and the cost classes your app declared it would accept.
Routing a run.
Your app reaches all of this over the endpoint or
/v1/runs.