Providers

Adding a provider is a row in a table — not a code change, not a release. Anything speaking the OpenAI API is already supported whether or not it is on the list below; the list exists so you do not have to go and find the base URL.

The Providers screen: backends grouped by api_base, each card showing model and backend counts, whether a credential is set, and how many are up

The Providers screen is where an endpoint and its credential are added and kept. A provider is a row in providers — since migration 0029 it is a record rather than a grouping the screen used to invent by bucketing backends on their api_base. Each card answers the question you actually have: how many models ride on it, whether a credential is set (never the credential itself), and how many of its models are up.

Add provider offers two ways in, which differ only in where the address comes from:

  • Cloud provider — pick a vendor from the catalogue, which carries every provider named on this page; the list has a filter, because eighty entries is more than anyone wants to scroll. Its wire protocol and the header it wants its key in are filled in for you, and its base URL where that is a fixed, verified address.

    Fifty-eight of the eighty-one carry a fixed address, each read off a source that dials it — LiteLLM's openai_compatible_endpoints and provider configs, go-ai-sdk, or the vendor's own documentation — and the catalogue's notes records which, so that when a vendor moves one there is somewhere to go and check.

    The other twenty-three carry a <placeholder>, and POST /admin/providers refuses to store an address with one still in it. Those are the cases nobody but you can fill in: a self-hosted engine runs wherever you started it, and an account-scoped endpoint encodes a resource, region, workspace or app that only your account knows (Azure, Bedrock, Vertex, Databricks, Snowflake, Cloudflare, Heroku). The entry still earns its place — it fills in the protocol and the header that vendor wants its key in, which is the half that is easy to get wrong.

  • Custom endpoint — type the address of anything else: a vLLM on the LAN, an Ollama on a workstation, a public vendor the catalogue does not list. This is a static provider whether the host is on your network or the internet: cloud means we preconfigured it, not that it is somewhere far away. Protocol defaults to openai, which is what almost everything speaks.

Nothing is served by adding a provider. It carries the endpoint and the credential; which of its models to expose is a separate, deliberate step on Provider models — because a cloud provider can front hundreds, and registering all of them is not what anyone means by adding one.

The catalogue

81 providers work today — 78 reached as-is, 3 through their own wire format. The count and these tables are checked against each other by tests/doc_claims.rs, so the number cannot drift away from the rows.

A caveat these tables are explicit about, because the number is otherwise a boast: "works" here means "is a configuration row that this proxy will forward to correctly", which follows from the endpoint being OpenAI-shaped. The ones exercised against real traffic in this repo's tests and on its dev cluster are marked ✓. The rest carry the base URL their vendor documents — check it against their docs before pasting it into production, because vendors move them and this file cannot notice.

reached as-is (OpenAI-compatible)
OpenRouter ✓ (fronts ~400 models)https://openrouter.ai/api/v1
OpenAI · Groq · DeepSeek · xAIapi.openai.com · api.groq.com · api.deepseek.com · api.x.ai
Together · Fireworks · Nebius · AtlasCloudfour endpoints, four rows
Mistral · Perplexity · Cerebras · SambaNovaapi.mistral.ai/v1 · api.perplexity.ai · api.cerebras.ai/v1 · api.sambanova.ai/v1
DeepInfra · Novita · Hyperbolic · Lambdafour endpoints, four rows
Z.ai · BigModel · Aliyun DashScope · Qwen Cloud
Moonshot / Kimi · Baidu Qianfan · AIHubMix
MiniMax · Volcengine Ark · Tencent Hunyuan · SarvamChinese and Indian clouds, same row shape
Baseten · Featherless · FriendliAI · Chutes
Nscale · GMI Cloud · Scaleway · OVHcloud
Cloudflare Workers AI · Vercel AI Gateway · v0 · Poe
NanoGPT · CometAPI · Inception · Morph
Clarifai · Weights & Biases · GradientAI · AI21
Snowflake Cortex · Anyscale · Heroku · CompactifAI
GitHub Models · GitHub Copilot
Amazon Bedrockhttps://bedrock-runtime.<region>.amazonaws.com/openai/v1, Bedrock API key as a bearer token
Coherehttps://api.cohere.ai/compatibility/v1
Google Vertex AIhttps://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/endpoints/openapi — see the API reference for the service-account credential
Azure OpenAI · Azure AIhttps://<resource>.openai.azure.com/openai/deployments/<deployment> with auth_header: api-key and auth_scheme: "" — the key goes in its own header with no Bearer prefix
NVIDIA NIM · Databricks · HuggingFace TGIintegrate.api.nvidia.com/v1 · a serving endpoint · any TGI /v1
vLLM ✓ · SGLang · llama.cpp ✓ · Ollamaself-hosted, same row shape
LM Studio · KoboldCpp · TabbyAPI · text-generation-webuilocal servers, same row shape
Xinference · Llamafile · Docker Model Runner · Lemonadelocal servers, same row shape
Voyage AI · Jina AI · Infinity · TEIembeddings and rerank — /v1/embeddings, /v1/rerank
reached through their own wire format
Anthropic"protocol": "anthropic" — Messages API, x-api-key, SSE re-framed to OpenAI chunks
Gemini"protocol": "gemini" — generateContent, model in the URL, x-goog-api-key
Z.ai (coding plan)"protocol": "anthropic" at https://api.z.ai/api/anthropic — catalogue key zai_coding_anthropic; an Anthropic-frontend client is passed through natively, which is what the plan's client fingerprinting requires. The endpoint has no /models, so the provider-create probe accepts its 200-wrapped 404 there and nowhere else. The plan's Codex door is separate: Responses API at https://api.z.ai/api/v1 (catalogue zai_coding_responses) — a /v1/responses request to an openai backend there is forwarded byte-exact, so name the frontend model after the upstream id and Codex traffic arrives as Codex sent it

Picking one

The Provider models screen offers a provider list: choose one and its base URL, protocol and auth header are filled in, so you paste a key and stop.

What is in that list is what this page documents an endpoint for. It names about a hundred providers and gives a host for thirty-odd of them; the rest are counted rather than specified, and seeding them would mean inventing base URLs. A list that confidently prefills a wrong endpoint is worse than one that admits it does not know.

So the list is a convenience, never a limit — anything speaking the OpenAI API works whether or not it is on it, and the address is always typeable. Two entries keep <region> placeholders (Bedrock, Vertex) rather than being prefilled with something that cannot resolve.

Adding one

A provider model is a name plus the providers that serve it, so there are two steps: the endpoint, then what you want off it. A model already registered somewhere can be attached to a second provider as well — that is what puts two hosts in one pool.

In the UI, on Providers, press Add provider and give it a credential. Then on Provider models:

The Provider models screen: each model with its provider, credential state, prices, cache TTL and context window

press Add model. The dialog asks for the provider first, then reads that endpoint's own answer to GET /v1/models and offers what it serves — not what a catalogue believes it offers — with the ones you have already registered marked. Pick one and the local name is filled in from it, editable before anything is created.

The order matters: naming a model first and finding it an address afterwards leaves a model that routes nowhere in between, and asks you to know the upstream name from memory before anything has offered it.

A provider that does not implement /v1/models says so in the dialog, and the upstream name stays typeable. Both writes are one intent: if attaching fails, the model created a moment earlier is removed rather than left behind as a name that routes nowhere and blocks the retry with a duplicate-name conflict.

By API, the same three calls:

# The endpoint and its key, once. A catalogue_key fills in the rest.
curl -sk -b /tmp/ck -X POST https://control:4001/admin/providers \
  -H 'content-type: application/json' \
  -d '{"catalogue_key":"openrouter","upstream_api_key":"sk-or-..."}'   # -> {"id":3}

curl -sk -b /tmp/ck -X POST https://control:4001/admin/provider-models \
  -H 'content-type: application/json' -d '{"name":"kimi-k2"}'          # -> {"id":7}

curl -sk -b /tmp/ck -X POST https://control:4001/admin/provider-models/7/backends \
  -H 'content-type: application/json' \
  -d '{"provider_id":3,"upstream_model":"moonshotai/kimi-k2"}'

Naming a provider_id settles the address, the protocol and the credential, so none of them may be sent alongside it — the API refuses rather than quietly preferring one source, which would leave you believing you had set a key here while the provider's is what actually gets sent.

Attaching by address still works, and still finds or creates a provider from it. That is how every backend was attached before providers were records, and it is what a LiteLLM import and every existing script does:

curl -sk -b /tmp/ck -X POST https://control:4001/admin/provider-models/7/backends \
  -H 'content-type: application/json' -d '{
    "api_base": "https://api.moonshot.ai/v1",
    "upstream_model": "moonshot-v1-128k",
    "upstream_api_key": "sk-..."
  }'

Pinned headers. extra_headers — on the backend attach, on POST /admin/providers, or on PATCH /admin/providers/{id} — is an object of name → value that every request to that provider carries, overriding the client's own header of the same name. The reason it exists: some endpoints judge a request by headers the caller used to supply, and user-agent is the one that comes up — Z.ai's coding plan refuses a key whose traffic does not look like Claude Code, and a pinned agent string is the part of that appearance the proxy can be told:

curl -sk -b /tmp/ck -X PATCH https://control:4001/admin/providers/3 \
  -H 'content-type: application/json' \
  -d '{"extra_headers": {"user-agent": "claude-cli/2.0.14 (external, cli)"}}'

A header that no client of the deployment sends is the honest case for this — it makes the proxy say what the operator decided, not fake what the client is. The body a non-Anthropic client sends still reads as what it is.

Posting that a second time against the same model with a different provider attaches it there too, and the two form one pool: the proxy chooses between them per request, which is what lets prefix-cache affinity send a conversation back to the machine that already has it warm. Two entries sharing a model_name in a LiteLLM config import to exactly that shape. Attaching the same provider twice is a 409: routing would treat it as two machines and send half the traffic to one it had already counted.

Prices, upstream_model and default_max_tokens live on the attachment rather than the model, because all three are facts about the provider — the same weights cost different amounts at different vendors, OpenRouter calls it google/gemini-2.5-flash where Google calls it gemini-2.5-flash, and only some providers demand a max_tokens. PATCH /admin/backends/{id} changes them; DELETE /admin/backends/{id} detaches one provider and leaves the model, its history and the provider itself alone.

Balancing between different models — failover, traffic splitting, spilling to the cloud — is a frontend model's job instead; see frontend models.

Credentials

upstream_api_key is encrypted at rest with FASTLLM_ENCRYPTION_KEY before it reaches Postgres, and the admin API never reads one back — the UI shows whether a credential is set, never what it is.

One credential per provider, however many models ride on it. That is the point of the split: rotating a key is one write, from the provider card's rotate key, or PATCH /admin/providers/{id} with a new upstream_api_key — not one write per model. An absent upstream_api_key there leaves the stored one alone, so renaming a provider does not require re-sending a key nothing can read back; "" clears it.

Two knobs exist because not every vendor puts the key in authorization: Bearer:

auth_headerthe header name. Default authorization
auth_schemethe prefix. Default Bearer; "" sends the key bare

Azure OpenAI is the case that needs both: auth_header: api-key and auth_scheme: "". Amazon Bedrock, despite the reputation, needs neither — its OpenAI-compatible endpoint takes a Bedrock API key as an ordinary bearer token, so it is a plain row like any other and there is no request signing.

Two providers speak their own language

Anthropic and Gemini do not expose an OpenAI-shaped endpoint, so they are reached through a translator rather than a base URL:

flowchart LR
    C["client<br/>OpenAI request"] --> P{"backend<br/>protocol?"}
    P -->|openai| B1["upstream<br/>bytes forwarded unchanged"]
    P -->|"anthropic +<br/>OpenAI client"| T1["translate →<br/>Messages API<br/>x-api-key"] --> B2["api.anthropic.com"]
    P -->|"anthropic +<br/>Anthropic client"| B4["upstream<br/>bytes forwarded unchanged<br/>(native passthrough)"]
    P -->|gemini| T2["translate →<br/>generateContent<br/>model in the URL"] --> B3["generativelanguage<br/>.googleapis.com"]
    B1 --> R1["response returned<br/>byte-for-byte"]
    B2 --> R2["SSE re-framed<br/>to OpenAI chunks"]
    B3 --> R2

Tool calling translates in both directions, streaming included, as do image and audio inputs. Translation is opt-in per backend: it costs parsing, and the whole latency argument for this gateway rests on not parsing. An openai backend's response body is never deserialised, which is why the two paths in that diagram are drawn differently — one forwards bytes, the other builds them.

The translation limits, field by field, are in the API reference.

Verified base URLs

Most providers are OpenAI-compatible, so they need no code at all — just a backend row pointing at their base URL. That includes OpenRouter, which itself fronts Anthropic, Gemini and several hundred other models in OpenAI format:

curl -X POST https://control/admin/provider-models/$MODEL_ID/backends \
  -H 'content-type: application/json' -b "$SESSION" \
  -d '{"api_base":"https://openrouter.ai/api/v1",
       "upstream_model":"anthropic/claude-sonnet-4",
       "upstream_api_key":"sk-or-..."}'

Verified base URLs for the OpenAI-compatible set:

providerapi_base
OpenRouterhttps://openrouter.ai/api/v1
OpenAIhttps://api.openai.com/v1
Groqhttps://api.groq.com/openai/v1
DeepSeekhttps://api.deepseek.com/v1
xAIhttps://api.x.ai/v1
Togetherhttps://api.together.xyz/v1
Fireworkshttps://api.fireworks.ai/inference/v1
Nebiushttps://api.studio.nebius.ai/v1
AtlasCloudhttps://api.atlascloud.ai/v1
AIHubMixhttps://aihubmix.com/v1
Z.aihttps://api.z.ai/api/paas/v4
BigModelhttps://open.bigmodel.cn/api/paas/v4
Aliyun DashScopehttps://dashscope.aliyuncs.com/compatible-mode/v1
Qwen Cloudhttps://dashscope-intl.aliyuncs.com/compatible-mode/v1
Moonshot / Kimihttps://api.moonshot.cn/v1, https://api.moonshot.ai/v1
Baidu Qianfanhttps://qianfan.baidubce.com/v2
GitHub Modelshttps://models.github.ai/inference
Ollamahttp://localhost:11434
Coherehttps://api.cohere.ai/compatibility/v1
Amazon Bedrockhttps://bedrock-runtime.<region>.amazonaws.com/openai/v1
Google Vertex AIhttps://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/endpoints/openapi

Bedrock needs no request signing. Its OpenAI-compatible endpoint takes a Bedrock API key as an ordinary bearer token, so it is a plain backend row like any other — create the key in the Bedrock console and put it in upstream_api_key.

What is deliberately absent

Deliberately absent, and not counted: providers whose API is not OpenAI-shaped and would need a fourth translator in src/protocol/ — Replicate, Predibase, Petals, Triton, WatsonX, OCI Generative AI, AWS SageMaker. Also absent are the non-LLM services a gateway has no business proxying blind: speech (Deepgram, ElevenLabs), image generation (Stability, Black Forest Labs, Recraft, Fal, RunwayML), vector stores (Milvus), and other people's gateways (Helicone, LiteLLM itself). Counting those would inflate the number without making anything work.

Where next

API and administrationVerified base URLs, and the per-field translation limits
What it can doRouting between providers, not just to them
OperationsWhere the encryption key lives, and why it cannot be regenerated