Architecture
Keep this current. Per CLAUDE.md, a change that adds a component, a role, an
endpoint crossing a plane boundary, or alters how the snapshot moves makes
these diagrams wrong, and redrawing them is part of that change's commit.
The two planes
One binary, three roles. The split between management and forwarding is a runtime flag, not a deployment boundary, so the same image is a single container in a lab and a scaled deployment in Kubernetes.
flowchart LR
client([OpenAI client])
aclient([Anthropic client<br/>Claude Code, SDKs])
subgraph dp["Data plane — --role proxy"]
msgs["/v1/messages<br/>Anthropic ⇄ OpenAI shape"]
auth[authenticate + authorise]
classify["classify prompt<br/>(only if classes exist)"]
route[resolve model, evaluate rules]
limit[rate limit + budget check]
fwd[forward opaque bytes]
xlate["translate<br/>(native protocols only)"]
snap[(Snapshot<br/>in memory)]
cache[(last-known-good<br/>on disk)]
end
subgraph cp["Control plane — --role control"]
admin["admin API + UI<br/>session auth, per-route permissions"]
fleetstore[("fleet health<br/>in memory, 30s TTL")]
build[build snapshot]
pg[(Postgres)]
end
backend([vLLM / SGLang / OpenRouter<br/>/ any OpenAI-compatible])
client -->|"Bearer sk-…"| auth --> route --> limit --> fwd --> backend
aclient -->|"x-api-key"| msgs -->|"chat completion"| auth
fwd -.->|"response translated back"| msgs
limit -.->|"backend.protocol ≠ openai"| xlate
xlate -.-> native([Anthropic / Gemini<br/>native API])
auth -.reads.-> snap
route -.reads.-> snap
limit -.reads.-> snap
admin --> pg
build --> pg
build -->|publishes| admin
snap <-->|"GET /snapshot (TLS, ETag)"| admin
snap -.writes.-> cache
cache -.restores on cold start.-> snap
fwd -->|"POST /usage (batched)"| admin
health["backend health<br/>probes + in-flight counts"] -->|"POST /health-report (every 10s)"| admin
admin --> fleetstore
fwd -.observed by.-> health
engines["engine load<br/>GET /metrics (every 2s)"] -.reads.-> route
--role all runs both boxes in one process; the snapshot is handed over in
memory instead of over HTTP, and the rate-limit reconciliation machinery is
inert because one process's counters are already global.
Why a snapshot, and why it is pre-flattened
The request path must do no I/O — the proxy's measured overhead against a real vLLM is zero, and a database round trip per request would cost more than the proxying itself. So every expensive question is answered once, when the snapshot is built, never per request:
flowchart TD
r[roles] --> p[permissions]
p --> w["wildcards expanded<br/>model/* → allow_all"]
w --> flat["Principal.allowed_models<br/>(a flat HashSet)"]
flat --> ask["request path:<br/>one set lookup"]
Per request that leaves: one SHA-256 of the bearer token, one hash lookup to a
principal, an expiry comparison, one set lookup for the model, and — when the
principal has a limit — an RwLock read plus up to two short mutex-guarded
bucket operations. No graph walk, no I/O, no lock held across an await, no
allocation beyond what the body already needed.
The Anthropic frontend
POST /v1/messages (src/protocol/messages.rs) faces the client rather than
the backend. The request is translated to a chat completion and handed to the
same proxy_request every other call uses, so routing, cache affinity,
budgets, rate limits and RBAC are not reimplemented and cannot drift; the
answer, streamed or not, is translated back to Anthropic's shape. x-api-key
is moved to Authorization before authentication.
When the routed backend itself speaks Anthropic, the translation stops at the
frontend: proxy_request holds the client's original bytes and forwards those
verbatim (model alias and auth aside), marked so the answer is handed back
without the OpenAI→Anthropic conversion. This is not an optimisation — a
twice-translated body has lost metadata, the per-block cache_control
markers and the client's key order, and that loss is exactly what an upstream
that fingerprints coding-agent traffic (Z.ai's coding plan, for one) detects
and refuses.
Known limits, because a translator should say what it does not do: thinking
blocks are dropped in both directions on the translated path (the native
passthrough has none), count_tokens is an estimate, and only
/chat/completions is translated — the other proxied suffixes are 501 on a
native-protocol backend.
Three execution modes
The dotted branch above is the whole of multi-provider support, and it is drawn dotted on purpose: it is not on the default path.
passthrough (protocol = openai) | native (anthropic backend behind /v1/messages) | translated (anthropic, gemini) | |
|---|---|---|---|
| request body | forwarded as-is, or one splice for a model alias | the client's own bytes, model alias aside | parsed and re-serialised into the native shape |
| response body | never parsed; forwarded byte for byte | forwarded byte for byte | parsed, re-framed into OpenAI chunks |
| usage | bounded tail buffer, one parse at end of stream | tail buffer plus a head-of-stream scanner | already parsed, exactly, during translation |
| endpoints | all seven proxied suffixes | /v1/messages only | /chat/completions only; the rest are 501 |
The native mode exists for the same fidelity reason as the passthrough — the
client's request reaches the provider exactly as the client built it — with
one addition the OpenAI case does not have: an Anthropic stream reports
input_tokens in its first event, which no bounded tail can hold, so
StreamUsage (src/protocol/anthropic.rs) mirrors the stream event-wise for
accounting while the bytes themselves go straight through.
| tool calling | passthrough, untouched | translated both directions, streaming included |
| image/audio input | passthrough, untouched | data: URLs translated inline; never fetched |
| overhead | zero measured against a real vLLM | one parse per frame |
Most providers are the left column, including OpenRouter — which is why
"support every provider genai supports" is mostly a configuration exercise
and not a code one. Only Anthropic and Gemini, addressed directly rather than
through an OpenAI-compatible gateway, are the right column.
The boundary is enforced, not merely intended: tests/native_protocols.rs
sends an intentionally odd-but-valid JSON document (unusual whitespace, key
order no serializer of ours would emit, a field we have no struct for) through
an openai backend and asserts the client receives those exact bytes. Any
accidental round trip through a parse shows up as a diff.
A request, end to end
sequenceDiagram
participant C as Client
participant P as Proxy
participant B as Backend
participant K as Control plane
C->>P: POST /v1/chat/completions
P->>P: SHA-256 → principal (401 if unknown/expired)
P->>P: classify prompt — only when classes are configured
P->>P: resolve model — frontend models evaluate rules, producing a fallback chain
P->>P: authorise the name the caller used (403 if ungranted)
P->>P: rate limit (429) and budget (402)
P->>P: translate request — only if backend.protocol ≠ openai
P->>B: forward, original bytes (or the translated ones)
B-->>P: 429/5xx → next backend, then the next model in the chain
B-->>P: response frames
P-->>C: same frames, never parsed
Note over P: tail buffer mirrors the last few KB
P->>P: at end of stream, parse once for usage
P-)K: POST /usage (batched, fire-and-forget)
K->>K: fold into budgets.tokens_used
P-)K: POST /health-report (every 10s, out of band)
Two decisions in that flow are load-bearing:
-
Authorisation is checked against the name the caller used. A request naming a frontend model is authorised against that frontend model; one naming a provider model directly is authorised against the provider model. A frontend model is how a model is exposed, so it is what gets granted — and a grant on it covers the chain it routes to, rather than being filtered per target. Adding a target to a frontend model therefore extends the reach of everyone holding it, which is acceptable only because editing one requires
config:write, itself all-or-nothing and already sufficient to grant any model outright. It is not a skeleton key: a provider model named directly still needs its own grant.This reversed the earlier rule, which required a grant on the resolved provider model. That rule pinned every grant to a provider model's name, so renaming one revoked access silently — migration 0029 did exactly that in production. See
.procoder/adr/0002-authorisation-moves-to-the-frontend-model.md.The "served here" check runs before authorisation, so an unknown model is a 404 for everyone and 403-vs-404 cannot be used to probe what exists.
-
Usage is read from a fixed-size tail buffer, parsed once at the end — never per frame. The response is still forwarded as opaque bytes. A translated response is the exception in the cheaper direction: its token counts were already parsed exactly, so it carries no tail buffer at all.
Administrative permissions
Admin routes are gated by a session and a per-route permission, drawn from
the same roles → permissions model the inference side uses: usage:read for
reads, key:create and key:revoke for key lifecycle, config:write for
everything else.
Two things an operator should know rather than discover:
config:writeis effectively administrative. A principal holding it can grant itself roles throughPOST /admin/principals/{id}/roles, sokey:create/key:revokeare a separation of duties, not a security boundary against it.- The
/admin/*404 for an unknown path is served outside the session gate, so an anonymous caller can tell which admin paths are not routes. It discloses no data, only the shape of the API.
Failure modes
| event | behaviour |
|---|---|
| control plane down, proxy warm | serves from memory; policy stops changing |
| control plane down, proxy cold | loads last-known-good from disk |
| cold start, no cache | starts, /health unhealthy, never crash-loops |
| snapshot invalid | keeps the previous one, logs once |
| key revoked | effective within the poll interval, ~1s |
| a model in a chain returns 429/5xx | the next model in the same rule serves it; nothing reached the client yet |
| every model in the chain refuses | the last upstream's own status and body are forwarded, not a synthetic 502 |
| one replica cannot reach a backend the rest can | the fleet's tally, returned in the reply to that replica's health report, contradicts it; the replica withdraws its own ejection and lets its next probe decide again |
| one replica genuinely is the only one that cannot | it hands the request to a sibling that can, one hop only; readiness cannot express "blind for one model" so the replica stays in rotation and this is what keeps it honest |
| Postgres down | control plane serves its last built snapshot; proxies unaffected |
| SIGTERM (a rollout) | stops accepting, lets in-flight generations finish, exits — up to --shutdown-grace (25s, under Kubernetes' 30s default) |
| usage report fails | dropped; never blocks a request |
| health report fails | dropped, logged at debug; GET /admin/fleet ages that replica out after 30s |
| upstream speaks an unexpected shape | translated backends only: the body fails rather than returning a plausible empty completion |
| snapshot names an unknown protocol | that backend is dropped with a logged reason, never silently treated as OpenAI |
Never crash-looping on a cold start is deliberate: under Kubernetes that would turn a control-plane outage into a data-plane outage, which is the failure this split exists to prevent.
Consistency, stated honestly
- Budgets are enforced after the fact. A request that blows the budget completes; the next is refused. Counting mid-stream would mean parsing every frame.
- Rate limits can overshoot by up to one reconciliation window during a sharp spike, because replicas enforce locally and reconcile periodically rather than sharing a counter on the request path.
- A replica with no recent traffic for a principal keeps a floor of
1/replicasof that principal's limit. Without it an idle replica's computed share collapses to zero and it refuses every request while the principal is far under budget — a worse failure than over-admitting. The floor bounds total allocation at under 2x the configured limit in the worst case (one busy replica, the rest idle), never more. - Policy changes propagate within one snapshot poll, not instantly.
- Semantic classification is deterministic and costs nothing when unused. With no prompt classes configured it is one atomic load and a length check. With classes, the fast tier is ~115µs of pure CPU; the refined tier is loaded only if some rule names a class that refines a fast-tier one, so a deployment that does not use it cannot pay for it. See semantic routing.
- Load balancing is an object, not a setting. A model pool is a named group of provider models with one policy; a rule points at it as a single target. That splits three questions that were previously tangled in one control: the order to try things in (a rule's targets), which of several models serves (a pool), and which of a model's own providers serves (the provider model). Each is set in exactly one place, and a pool is reusable across rules — which a policy attached to one rule's target list could never be.
- Every routing rule is terminal. A rule that matches decides everything
about the request — route it, refuse it, or delegate the whole decision to
another frontend model's chain. Firewalls have non-terminating rules that
mark and continue; that is deliberately not copied, because "first match
wins and the matching rule decided" is what lets
/admin/routing/dry-runanswer with one rule name rather than a trace. Modifiers (a usagetag, a target-selectionpolicy) are fields on a routing rule, never separate passes. - Two routing conditions are deliberately non-deterministic.
max_inflight_per_backendreads live in-flight counters — the engine's own, scraped from its Prometheus/metricsby each proxy in the background, so the ceiling means the same thing however many proxies are running, falling back to this replica's count for a backend that publishes none (which is detected by asking, not configured) — and the time-window conditions read the clock, so identical requests can route differently and prefix affinity stops applying to the traffic they divert. Every other condition is a pure function of the request. This is the same opt-in-visibly line the passthrough/translate split draws. - The queue forms in the proxy, not in the engine (flow control, set per
backend on the
model_backendsrow; seedocs/operations/configuration.md). vLLM accepts everything and queues the surplus itself, unbounded, which makes every request slow rather than failing any. A backend with a gate carries it in the snapshot like any other backend setting; the proxy's engine scraper halves the gate's ceiling while the engine'snum_requests_waitingsits at or above its high-water mark, and grows it by one per scrape once the queue drains. Requests past the ceiling wait in the proxy, bounded, and are refused with 503 andRetry-Afterbeyond that. The request path only awaits a semaphore, so it still performs no I/O. Each replica holds its own ceiling -- sharing it would mean a network round trip per request -- and the engine's queue depth, which already reflects every replica's traffic, is what keeps them honest. Each gate's state rides the existing health report to the control plane, which is how the Fleet and Models pages show it.
Behaviour notes
- Retries only happen before any byte has been forwarded. Once the response is committed a mid-stream failure propagates as-is — it cannot be silently retried without corrupting the stream.
- 5xx is retried, 4xx is not. A client error retried across every node is the same client error three times.
- The last backend's response is forwarded verbatim. A 5xx is only retried while another backend remains; when none does, the upstream's own status and body reach the client rather than a synthetic 502. On a single-node pool that means every error keeps the engine's diagnostics.
- Audio endpoints take
multipart/form-data.modelis read from the form field and the upload is forwarded byte for byte, content-type and boundary intact. An alias splices the new name into that one field rather than re-encoding the body. https://backends work, so a TLS-terminated or hosted endpoint can sit in the same config as cluster-local nodes. System root certificates are used, falling back to the bundled Mozilla set.- When every backend in a model is unhealthy, the request skips to the next model or the deployment-wide fallback rather than sending to a known-failed node.
tests/failover.rsverifies the end-to-end behaviour. - The client's
Authorizationheader is never forwarded. It authenticates the client to the proxy; the upstream gets the backend's own key or none — in whichever header that provider reads it from. - An OpenAI-compatible backend's response is never parsed. Bytes are forwarded verbatim, which is why proxied overhead measures at zero;
tests/native_protocols.rspins it against an intentionally odd-but-valid payload. Only a backend explicitly configured for a native protocol is translated, and only there is a response body read. - A backend's identity covers its whole configuration. Rotating an upstream key, or changing a backend's protocol, produces a new routing entry rather than reusing the live one — otherwise a reload would keep serving with the old credential, since backend objects are carried across reloads to preserve their in-flight counts.
- Affinity keys hash the raw request prefix, not parsed fields. JSON does not guarantee field order, but order is stable per client, which is all affinity needs — a client that reorders per request degrades to least-loaded rather than misrouting.
- A rate-limited request gets
429withRetry-After, checked after authorisation and model resolution but before the request is dispatched upstream — nothing is forwarded on a rejected request. See "Rate limits" above.