One HTTP request that returns a finished agentic answer, executed on someone's own Claude Code or Codex, with tools your app defines and executes.
JavaScript
import { Soba } from "@soba-so/sdk";
const soba = new Soba({ apiKey: process.env.SOBA_KEY });
const run = await soba.run({
user: session.user.id,
model: "auto",
prompt: "Summarise everything I saved this week",
tools: [listNotes, readNote], // your code, executed in your app
});
for await (const event of run) {
if (event.type === "delta") process.stdout.write(event.text);
if (event.type === "done") console.log(event.usage);
}
A run is one request that comes back finished. It reaches everything the endpoint
reaches and, unlike the endpoint, the user's own Claude Code or Codex on their
subscription, and it streams back the
normalised event stream.
Why this exists alongside the endpoint#
chat/completions is stateless, and the caller owns the loop:
you send messages, you get a tool call back, you execute it, you send everything
again. That works for a model. It cannot work for an agent.
Claude Code and Codex run their own loop, and they never hand a tool call back to wait
on an HTTP round trip. So reaching them means Soba owns the loop, and your tools
have to be registered out of band rather than returned mid-conversation.
That is the whole difference between the two entry points:
|
chat/completions |
soba.run() |
| Who owns the loop |
You |
Soba |
| Tools |
Returned to you, per turn |
Registered up front, called back |
| Reaches model runtimes |
yes |
yes |
| Reaches Claude Code and Codex |
no |
yes |
| Shape |
The request you already send |
One request, one finished answer |
The model calls them wherever the run is executing, the call comes back to your app, and
your code executes it. Nothing about your tools moves to Soba: they run where they
already run, against the data they already touch. See
App-defined tools.
Over HTTP that is a frame on the run's stream and a request back the other way:
JSON
{"type":"tool_call","runId":"r_01","callId":"c_1","name":"lookup_invoices","input":{"customer":"Acme"}}
Terminal
curl https://api.soba.so/v1/runs/r_01/tool_result \
-H "Authorization: Bearer $SOBA_KEY" \
-H "Content-Type: application/json" \
-d '{ "callId": "c_1", "result": { "open": 3, "total": "$14,200" } }'
The SDK does that for you: your execute returns, and the result is on its way.
isError: true marks a failure the model should see and can react to, rather than
one that kills the run.
Tools can also be registered from an MCP server you already operate, which is what
lets a run reach a toolset you did not write inline.
Over raw HTTP#
Terminal
curl https://api.soba.so/v1/runs \
-H "Authorization: Bearer $SOBA_KEY" \
-H "Content-Type: application/json" \
-d '{
"user": "usr_123",
"model": "auto",
"prompt": "Summarise everything I saved this week",
"route": { "prefer": ["user-hardware", "user-subscription"] }
}'
POST /v1/runs takes the same shape the SDK sends and streams the same events back.
The SDK is a convenience over it, not a different product. The run's id arrives in a
Soba-Run-Id header, and three routes
(tool_result, approval, cancel) answer a run that is
still in flight.
The events#
JSON
{"type":"status","message":"Starting","tier":"user-subscription"}
{"type":"delta","text":"You saved eleven things this week. "}
{"type":"done","usage":{"model":"sonnet","costClass":"user-subscription","inputTokens":1840,"outputTokens":210,"cacheReadTokens":0,"cacheWriteTokens":0,"webSearches":0}}
Five event types. delta appends, thinking replaces, status carries the
tier that served the run before any output, done carries what it consumed, and
error is the other way a run ends.
Full shapes on The event stream.
Declaring what a run may use#
JavaScript
const run = await soba.run({
user: session.user.id,
prompt: "…",
route: {
prefer: ["user-hardware", "user-subscription"],
allow: ["user-hardware", "user-subscription", "app"],
maxCostMicros: 250000,
},
});
You declare the cost classes you accept, not a runtime. allow is a hard ceiling and is
enforced on the machine as well as by Soba. See Routing a run.
Most apps never write route by hand: a plan compiles to one on every
run, so pricing and routing stay the same decision.
Serving the run yourself#
A user with nothing reachable is served by your app, on your own key, the way you serve
them today. Give the SDK a fallback and it calls it in your process whenever a run is
handed back:
JavaScript
const soba = new Soba({
apiKey: process.env.SOBA_KEY,
// Runs in your server, with your key, only when the user's own compute cannot serve the run.
fallback: async (run) =>
openai.chat.completions.create({
model: "gpt-5",
messages: run.messages,
stream: true,
stream_options: { include_usage: true },
}),
});
The fallback receives what the run was asked to do: the messages, the system prompt,
your tools and any maxCostMicros, which only your code can enforce here. Return a
stream or a finished completion; the SDK streams it back as the run's
events, with tier: "app" on the opening status, and
reports its usage to Soba when it ends. Your key never leaves your server.