Some runtimes run their own agent loop and some are a model with nobody driving. The difference decides which entry point you use and whether policy is advisory or enforced.
Runtimes are not fungible. Claude Code, Codex and a raw provider SDK differ in
tools, session semantics, permission models and cost. Pretending otherwise is how a
routing layer leaks.
So instead of a lowest-common-denominator API, a machine declares what it can
serve and an app declares what it will accept, and Soba matches the two.
Two families#
Agent runtimes (Claude Code and Codex) run their own loop, tools and approvals.
Soba normalises what comes out of them, and nothing else.
Model runtimes (an OpenAI-compatible endpoint, Ollama) are completion plus
tool-calling, with nobody driving. Soba supplies the loop.
The split has two consequences worth stating plainly:
|
Agent runtime |
Model runtime |
| Who runs the loop |
The CLI |
Soba |
| Where policy is enforced |
The CLI is given the permitted tool set and trusted to honour it |
In-process, immediately before the tool runs |
| Strength |
Advisory |
Enforced: case-folded, denylist wins over everything |
Reachable from chat/completions |
no |
yes |
Reachable from soba.run() |
yes |
yes |
That last pair is the practical one. The endpoint is stateless and the
caller owns the loop, so it reaches models. Reaching someone's Claude Code or Codex
means letting Soba own the loop, which is soba.run().
The loop takes its toolset by injection rather than importing one. On the user's
machine it gets the local filesystem and shell; served directly by Soba it gets
neither, because there is no user filesystem there.
Same loop, two capability profiles, and the difference between "runs on your laptop"
and "runs in a datacentre with your token" stays visible instead of hiding behind a
flag.
What a machine advertises#
For each runtime it can serve: an id, a display name, a version, the models available,
the native tool names, its cost class, pricing if the class is a
metered one, whether it can execute tool schemas your app supplies, and whether it is
signed in.
A runtime that is verifiably signed out is withheld rather than advertised, so it
cannot be routed to and fail at spawn. The owner is told, with the command that fixes
it. See Cost classes.
Which runtimes are supported#
| Runtime |
|
claude-code |
Agent. The user's Claude Pro or Max plan |
codex |
Agent. The user's ChatGPT Plus or Pro plan |
anthropic-api, openai-api |
Model. The machine owner's own provider key (user-key) |
| Ollama |
Model. user-hardware: free because the silicon is already bought |
| Any OpenAI-compatible endpoint |
Model. Configuration, not code |
Ollama is detected by probing its port rather than looking on PATH, and reports the
models you have actually pulled. Because it runs on the machine it is also
capability-complete: it gets the filesystem and the shell.
The Soba app lists what a machine is offering, and on a headless box
soba-worker --status prints the same report — including the models a runtime can be
asked for and which one it uses when a run names none. Set that one on the runtime's
page, or in defaultModels; see Which model a run gets.
That applies to the agents too, not only the endpoints. Claude Code accepts haiku,
sonnet, opus and fable; pick one on its page and a run that names no model lands
on it instead of on whatever the CLI's own settings say. Leave it on Auto and the
CLI decides, which is the behaviour every machine had before the choice existed.
Small models are unreliable at tool calling
A model that advertises tool support may still emit tool calls as prose, or call the
wrong tool and then invent the result. Ollama's advertised capabilities do not tell you
which models are safe here. Choose the model deliberately, and expect small ones to be
unreliable at this rather than broken at it.
Text a model produces is never promoted into a real tool call, however much it looks
like one. Text in a response can come from a file or a page the model just read, so
promoting it would turn any document into a way to steer execution.