Ollama, Together, Groq, vLLM and anything self-hosted speak the same API, so adding one is configuration rather than code, and the cost class is part of that configuration.
~/.soba/models.json:
JSON
[
{ "id": "together", "baseUrl": "https://api.together.xyz/v1",
"apiKeyEnv": "TOGETHER_API_KEY", "costClass": "user-key",
"pricing": { "inputMicrosPerMTok": 150000, "outputMicrosPerMTok": 600000 } }
]
Ollama, OpenRelay, Together, Groq, DeepInfra, Fireworks, vLLM and anything
self-hosted all speak the OpenAI-compatible API, so there is no adapter to write.
Fields#
| Field |
|
id |
The runtime id this endpoint is advertised under |
baseUrl |
The OpenAI-compatible base, including /v1 |
apiKeyEnv |
Preferred. The name of an environment variable holding the key |
apiKey |
The key inline. Works, but see below |
costClass |
Required: user-hardware or user-key. No default; see below |
pricing |
inputMicrosPerMTok / outputMicrosPerMTok. Required for metered classes |
Prefer apiKeyEnv over apiKey#
A key written into a config file outlives the reason it was put there and gets
copied around with it.
A configured key that isn't set is treated exactly like a signed-out CLI: the runtime
is withheld rather than advertised, and reported to you with the fix, rather than failing
every run.
Put the key in ~/.soba/service.env (mode 0600) so it survives a service install
without appearing in a world-readable unit file. See
Keeping it running.
There is no default cost class#
A row without a valid one is skipped.
Guessing who pays is the mistake the field exists to prevent. Declaring it is a claim
about whose money this is, and the only money reachable from this file is the
machine owner's. Nothing here can spend yours, which is why the file is
trusted to say it, and why the only classes it accepts are user-hardware and
user-key.
Prices are integer micros#
Throughout: inputMicrosPerMTok is micros per million input tokens. No floats, no
currency strings.
route.maxCostMicros stops a run that reaches its ceiling, and that check happens
between turns, because nobody knows a turn's cost before it happens. The
guarantee is "stops as soon as it knows", not "never exceeds".
Ollama needs no entry#
Ollama is detected by probing its port, and reports the models you have actually
pulled. It is user-hardware, so it carries no pricing. Add an entry only when you
want to point at a non-default host. SOBA_OLLAMA_URL also works.
Nothing here downloads a model. ollama pull does, and the next probe finds it —
the list an endpoint advertises is the list it answered with, never a guess.
If Ollama is running but has nothing pulled, the machine reports it as present and
unusable rather than missing, with the one command that fixes it. "Not installed"
and "installed with nothing in it" are different problems.
Which model a run gets#
A run may name one. If it does not — or names one this machine has not pulled — the
model comes from defaultModels in ~/.soba/policy.json, keyed by runtime id:
JSON
{ "defaultModels": { "ollama": "qwen2.5-coder:7b" } }
The Soba app writes this for you: open the runtime's page, go to Models and press
one. Without it the answer is whichever model the endpoint happened to list first,
which is nobody's decision.
The same key holds an agent CLI's choice, where the value is one of the aliases it
accepts rather than a model it has pulled — { "defaultModels": { "claude-code": "opus" } }.
Removing the key, which the Auto row on that page does, puts the decision back
where it was: with the CLI's own settings.
A run that names a model you do not have still runs, on the default — and the app is
told which model answered, in a status event. It used to substitute in silence,
which turned "you never pulled that model" into "the agent behaved strangely" with
nothing connecting the two.
Choose the model deliberately
Small models are unreliable at tool calling, and Ollama's advertised capabilities
do not tell you which. See Agents and models.