Architecture

Keep this current. Per CLAUDE.md, a change that adds a component, a role, an endpoint crossing a plane boundary, or alters how the snapshot moves makes these diagrams wrong, and redrawing them is part of that change's commit.

The two planes

One binary, three roles. The split between management and forwarding is a runtime flag, not a deployment boundary, so the same image is a single container in a lab and a scaled deployment in Kubernetes.

flowchart LR
    client([OpenAI client])
    aclient([Anthropic client<br/>Claude Code, SDKs])

    subgraph dp["Data plane — --role proxy"]
        msgs["/v1/messages<br/>Anthropic ⇄ OpenAI shape"]
        auth[authenticate + authorise]
        classify["classify prompt<br/>(only if classes exist)"]
        route[resolve model, evaluate rules]
        limit[rate limit + budget check]
        fwd[forward opaque bytes]
        xlate["translate<br/>(native protocols only)"]
        snap[(Snapshot<br/>in memory)]
        cache[(last-known-good<br/>on disk)]
    end

    subgraph cp["Control plane — --role control"]
        admin["admin API + UI<br/>session auth, per-route permissions"]
        fleetstore[("fleet health<br/>in memory, 30s TTL")]
        build[build snapshot]
        pg[(Postgres)]
    end

    backend([vLLM / SGLang / OpenRouter<br/>/ any OpenAI-compatible])

    client -->|"Bearer sk-…"| auth --> route --> limit --> fwd --> backend
    aclient -->|"x-api-key"| msgs -->|"chat completion"| auth
    fwd -.->|"response translated back"| msgs
    limit -.->|"backend.protocol ≠ openai"| xlate
    xlate -.-> native([Anthropic / Gemini<br/>native API])
    auth -.reads.-> snap
    route -.reads.-> snap
    limit -.reads.-> snap

    admin --> pg
    build --> pg
    build -->|publishes| admin
    snap <-->|"GET /snapshot (TLS, ETag)"| admin
    snap -.writes.-> cache
    cache -.restores on cold start.-> snap
    fwd -->|"POST /usage (batched)"| admin
    health["backend health<br/>probes + in-flight counts"] -->|"POST /health-report (every 10s)"| admin
    admin --> fleetstore
    fwd -.observed by.-> health
    engines["engine load<br/>GET /metrics (every 2s)"] -.reads.-> route

--role all runs both boxes in one process; the snapshot is handed over in memory instead of over HTTP, and the rate-limit reconciliation machinery is inert because one process's counters are already global.

Why a snapshot, and why it is pre-flattened

The request path must do no I/O — the proxy's measured overhead against a real vLLM is zero, and a database round trip per request would cost more than the proxying itself. So every expensive question is answered once, when the snapshot is built, never per request:

flowchart TD
    r[roles] --> p[permissions]
    p --> w["wildcards expanded<br/>model/* → allow_all"]
    w --> flat["Principal.allowed_models<br/>(a flat HashSet)"]
    flat --> ask["request path:<br/>one set lookup"]

Per request that leaves: one SHA-256 of the bearer token, one hash lookup to a principal, an expiry comparison, one set lookup for the model, and — when the principal has a limit — an RwLock read plus up to two short mutex-guarded bucket operations. No graph walk, no I/O, no lock held across an await, no allocation beyond what the body already needed.

The Anthropic frontend

POST /v1/messages (src/protocol/messages.rs) faces the client rather than the backend. The request is translated to a chat completion and handed to the same proxy_request every other call uses, so routing, cache affinity, budgets, rate limits and RBAC are not reimplemented and cannot drift; the answer, streamed or not, is translated back to Anthropic's shape. x-api-key is moved to Authorization before authentication.

When the routed backend itself speaks Anthropic, the translation stops at the frontend: proxy_request holds the client's original bytes and forwards those verbatim (model alias and auth aside), marked so the answer is handed back without the OpenAI→Anthropic conversion. This is not an optimisation — a twice-translated body has lost metadata, the per-block cache_control markers and the client's key order, and that loss is exactly what an upstream that fingerprints coding-agent traffic (Z.ai's coding plan, for one) detects and refuses.

Known limits, because a translator should say what it does not do: thinking blocks are dropped in both directions on the translated path (the native passthrough has none), count_tokens is an estimate, and only /chat/completions is translated — the other proxied suffixes are 501 on a native-protocol backend.

Three execution modes

The dotted branch above is the whole of multi-provider support, and it is drawn dotted on purpose: it is not on the default path.

passthrough (protocol = openai)native (anthropic backend behind /v1/messages)translated (anthropic, gemini)
request bodyforwarded as-is, or one splice for a model aliasthe client's own bytes, model alias asideparsed and re-serialised into the native shape
response bodynever parsed; forwarded byte for byteforwarded byte for byteparsed, re-framed into OpenAI chunks
usagebounded tail buffer, one parse at end of streamtail buffer plus a head-of-stream scanneralready parsed, exactly, during translation
endpointsall seven proxied suffixes/v1/messages only/chat/completions only; the rest are 501

The native mode exists for the same fidelity reason as the passthrough — the client's request reaches the provider exactly as the client built it — with one addition the OpenAI case does not have: an Anthropic stream reports input_tokens in its first event, which no bounded tail can hold, so StreamUsage (src/protocol/anthropic.rs) mirrors the stream event-wise for accounting while the bytes themselves go straight through. | tool calling | passthrough, untouched | translated both directions, streaming included | | image/audio input | passthrough, untouched | data: URLs translated inline; never fetched | | overhead | zero measured against a real vLLM | one parse per frame |

Most providers are the left column, including OpenRouter — which is why "support every provider genai supports" is mostly a configuration exercise and not a code one. Only Anthropic and Gemini, addressed directly rather than through an OpenAI-compatible gateway, are the right column.

The boundary is enforced, not merely intended: tests/native_protocols.rs sends an intentionally odd-but-valid JSON document (unusual whitespace, key order no serializer of ours would emit, a field we have no struct for) through an openai backend and asserts the client receives those exact bytes. Any accidental round trip through a parse shows up as a diff.

A request, end to end

sequenceDiagram
    participant C as Client
    participant P as Proxy
    participant B as Backend
    participant K as Control plane

    C->>P: POST /v1/chat/completions
    P->>P: SHA-256 → principal (401 if unknown/expired)
    P->>P: classify prompt — only when classes are configured
    P->>P: resolve model — frontend models evaluate rules, producing a fallback chain
    P->>P: authorise the name the caller used (403 if ungranted)
    P->>P: rate limit (429) and budget (402)
    P->>P: translate request — only if backend.protocol ≠ openai
    P->>B: forward, original bytes (or the translated ones)
    B-->>P: 429/5xx → next backend, then the next model in the chain
    B-->>P: response frames
    P-->>C: same frames, never parsed
    Note over P: tail buffer mirrors the last few KB
    P->>P: at end of stream, parse once for usage
    P-)K: POST /usage (batched, fire-and-forget)
    K->>K: fold into budgets.tokens_used
    P-)K: POST /health-report (every 10s, out of band)

Two decisions in that flow are load-bearing:

  • Authorisation is checked against the name the caller used. A request naming a frontend model is authorised against that frontend model; one naming a provider model directly is authorised against the provider model. A frontend model is how a model is exposed, so it is what gets granted — and a grant on it covers the chain it routes to, rather than being filtered per target. Adding a target to a frontend model therefore extends the reach of everyone holding it, which is acceptable only because editing one requires config:write, itself all-or-nothing and already sufficient to grant any model outright. It is not a skeleton key: a provider model named directly still needs its own grant.

    This reversed the earlier rule, which required a grant on the resolved provider model. That rule pinned every grant to a provider model's name, so renaming one revoked access silently — migration 0029 did exactly that in production. See .procoder/adr/0002-authorisation-moves-to-the-frontend-model.md.

    The "served here" check runs before authorisation, so an unknown model is a 404 for everyone and 403-vs-404 cannot be used to probe what exists.

  • Usage is read from a fixed-size tail buffer, parsed once at the end — never per frame. The response is still forwarded as opaque bytes. A translated response is the exception in the cheaper direction: its token counts were already parsed exactly, so it carries no tail buffer at all.

Administrative permissions

Admin routes are gated by a session and a per-route permission, drawn from the same roles → permissions model the inference side uses: usage:read for reads, key:create and key:revoke for key lifecycle, config:write for everything else.

Two things an operator should know rather than discover:

  • config:write is effectively administrative. A principal holding it can grant itself roles through POST /admin/principals/{id}/roles, so key:create/key:revoke are a separation of duties, not a security boundary against it.
  • The /admin/* 404 for an unknown path is served outside the session gate, so an anonymous caller can tell which admin paths are not routes. It discloses no data, only the shape of the API.

Failure modes

eventbehaviour
control plane down, proxy warmserves from memory; policy stops changing
control plane down, proxy coldloads last-known-good from disk
cold start, no cachestarts, /health unhealthy, never crash-loops
snapshot invalidkeeps the previous one, logs once
key revokedeffective within the poll interval, ~1s
a model in a chain returns 429/5xxthe next model in the same rule serves it; nothing reached the client yet
every model in the chain refusesthe last upstream's own status and body are forwarded, not a synthetic 502
one replica cannot reach a backend the rest canthe fleet's tally, returned in the reply to that replica's health report, contradicts it; the replica withdraws its own ejection and lets its next probe decide again
one replica genuinely is the only one that cannotit hands the request to a sibling that can, one hop only; readiness cannot express "blind for one model" so the replica stays in rotation and this is what keeps it honest
Postgres downcontrol plane serves its last built snapshot; proxies unaffected
SIGTERM (a rollout)stops accepting, lets in-flight generations finish, exits — up to --shutdown-grace (25s, under Kubernetes' 30s default)
usage report failsdropped; never blocks a request
health report failsdropped, logged at debug; GET /admin/fleet ages that replica out after 30s
upstream speaks an unexpected shapetranslated backends only: the body fails rather than returning a plausible empty completion
snapshot names an unknown protocolthat backend is dropped with a logged reason, never silently treated as OpenAI

Never crash-looping on a cold start is deliberate: under Kubernetes that would turn a control-plane outage into a data-plane outage, which is the failure this split exists to prevent.

Consistency, stated honestly

  • Budgets are enforced after the fact. A request that blows the budget completes; the next is refused. Counting mid-stream would mean parsing every frame.
  • Rate limits can overshoot by up to one reconciliation window during a sharp spike, because replicas enforce locally and reconcile periodically rather than sharing a counter on the request path.
  • A replica with no recent traffic for a principal keeps a floor of 1/replicas of that principal's limit. Without it an idle replica's computed share collapses to zero and it refuses every request while the principal is far under budget — a worse failure than over-admitting. The floor bounds total allocation at under 2x the configured limit in the worst case (one busy replica, the rest idle), never more.
  • Policy changes propagate within one snapshot poll, not instantly.
  • Semantic classification is deterministic and costs nothing when unused. With no prompt classes configured it is one atomic load and a length check. With classes, the fast tier is ~115µs of pure CPU; the refined tier is loaded only if some rule names a class that refines a fast-tier one, so a deployment that does not use it cannot pay for it. See semantic routing.
  • Load balancing is an object, not a setting. A model pool is a named group of provider models with one policy; a rule points at it as a single target. That splits three questions that were previously tangled in one control: the order to try things in (a rule's targets), which of several models serves (a pool), and which of a model's own providers serves (the provider model). Each is set in exactly one place, and a pool is reusable across rules — which a policy attached to one rule's target list could never be.
  • Every routing rule is terminal. A rule that matches decides everything about the request — route it, refuse it, or delegate the whole decision to another frontend model's chain. Firewalls have non-terminating rules that mark and continue; that is deliberately not copied, because "first match wins and the matching rule decided" is what lets /admin/routing/dry-run answer with one rule name rather than a trace. Modifiers (a usage tag, a target-selection policy) are fields on a routing rule, never separate passes.
  • Two routing conditions are deliberately non-deterministic. max_inflight_per_backend reads live in-flight counters — the engine's own, scraped from its Prometheus /metrics by each proxy in the background, so the ceiling means the same thing however many proxies are running, falling back to this replica's count for a backend that publishes none (which is detected by asking, not configured) — and the time-window conditions read the clock, so identical requests can route differently and prefix affinity stops applying to the traffic they divert. Every other condition is a pure function of the request. This is the same opt-in-visibly line the passthrough/translate split draws.
  • The queue forms in the proxy, not in the engine (flow control, set per backend on the model_backends row; see docs/operations/configuration.md). vLLM accepts everything and queues the surplus itself, unbounded, which makes every request slow rather than failing any. A backend with a gate carries it in the snapshot like any other backend setting; the proxy's engine scraper halves the gate's ceiling while the engine's num_requests_waiting sits at or above its high-water mark, and grows it by one per scrape once the queue drains. Requests past the ceiling wait in the proxy, bounded, and are refused with 503 and Retry-After beyond that. The request path only awaits a semaphore, so it still performs no I/O. Each replica holds its own ceiling -- sharing it would mean a network round trip per request -- and the engine's queue depth, which already reflects every replica's traffic, is what keeps them honest. Each gate's state rides the existing health report to the control plane, which is how the Fleet and Models pages show it.

Behaviour notes

  • Retries only happen before any byte has been forwarded. Once the response is committed a mid-stream failure propagates as-is — it cannot be silently retried without corrupting the stream.
  • 5xx is retried, 4xx is not. A client error retried across every node is the same client error three times.
  • The last backend's response is forwarded verbatim. A 5xx is only retried while another backend remains; when none does, the upstream's own status and body reach the client rather than a synthetic 502. On a single-node pool that means every error keeps the engine's diagnostics.
  • Audio endpoints take multipart/form-data. model is read from the form field and the upload is forwarded byte for byte, content-type and boundary intact. An alias splices the new name into that one field rather than re-encoding the body.
  • https:// backends work, so a TLS-terminated or hosted endpoint can sit in the same config as cluster-local nodes. System root certificates are used, falling back to the bundled Mozilla set.
  • When every backend in a model is unhealthy, the request skips to the next model or the deployment-wide fallback rather than sending to a known-failed node. tests/failover.rs verifies the end-to-end behaviour.
  • The client's Authorization header is never forwarded. It authenticates the client to the proxy; the upstream gets the backend's own key or none — in whichever header that provider reads it from.
  • An OpenAI-compatible backend's response is never parsed. Bytes are forwarded verbatim, which is why proxied overhead measures at zero; tests/native_protocols.rs pins it against an intentionally odd-but-valid payload. Only a backend explicitly configured for a native protocol is translated, and only there is a response body read.
  • A backend's identity covers its whole configuration. Rotating an upstream key, or changing a backend's protocol, produces a new routing entry rather than reusing the live one — otherwise a reload would keep serving with the old credential, since backend objects are carried across reloads to preserve their in-flight counts.
  • Affinity keys hash the raw request prefix, not parsed fields. JSON does not guarantee field order, but order is stable per client, which is all affinity needs — a client that reorders per request degrades to least-loaded rather than misrouting.
  • A rate-limited request gets 429 with Retry-After, checked after authorisation and model resolution but before the request is dispatched upstream — nothing is forwarded on a rejected request. See "Rate limits" above.