Changelog
Notable changes, newest first. Format follows Keep a Changelog.
Commit bodies carry the reasoning and the measurements and remain the better source for why anything is the way it is; this file is the summary.
Unreleased
Added
- Load balancing is per backend model, not per process.
--policywas a deployment-wide flag, which is the wrong shape the moment one control plane serves both kinds of pool — two identical local replicas sharing a prefix cache wantcache-affinity, three hosted providers of differing speed wantlowest-latency, and a flag can only be one of them. Each backend model may now carry its own (migration 0028,policyonPOST/PATCH /admin/models, a control on the Backend models screen). Unset means the deployment default, so an existing database behaves exactly as it did. - The price sync can replace a price that is already set. It never
overwrote by design — a negotiated rate must not be replaced by a list
price — but that left a model priced wrongly unreachable from the UI,
including one sitting at
0, which reads as free. The preview now has a "replace prices that are already set" toggle, off by default, that re-previews as it changes.
Changed
- Two words for two things: backend model and frontend model. A backend
model is what a request is routed to (one name, its backends, its
load-balancing policy); a frontend model is what a client asks for (rules
and weights resolving to a chain of backend models). The UI, the navigation
and the documentation use them consistently; the admin API still spells them
modelsandvirtual-modelsin its paths, so every existing script and the OpenAPI description keep working.
Fixed
-
Declared context windows never reached routing.
Registrycarried acontext_lengthmap whose doc comment said it was filled from the snapshot, and nothing ever filled it — sorouting::candidates' context-window fallback, which demotes a model whose window provably cannot hold the request, could not fire in any production build. The column, the admin API field and the routing code were all present and correct; only the wiring between them was missing. -
The Kubernetes operator earns its keep. It was removed earlier in this cycle for reconciling two Deployments a chart already produces; it is back because the four things a chart genuinely cannot do are now implemented and verified against a live cluster:
- Ordered upgrades. The two planes share a database schema, so
spec.imagerolls the control plane first and holds the gateway at the image it is running until that has finished. Verified with a deliberately unpullable tag: the control plane went down, the gateway kept serving on the old image, and theUpgradingcondition said which and why. - Rotation that takes effect. Both pod templates carry a hash of the resolved Secret material, so rotating the proxy token — or cert-manager renewing the control-plane certificate — rolls the pods instead of silently doing nothing until an unrelated restart.
- Preflight. Every referenced Secret is resolved and checked before
anything is applied; a missing key or a short encryption key becomes a
condition naming the Secret and the key rather than pods in
CreateContainerConfigError.encryptionKeyis immutable, enforced by the API server through a CEL rule. - A finished install.
bootstraprunsset-passwordas a Job once the control plane is ready, so the deployment ends with a UI that can be signed into. Verified end to end:POST /loginreturns 200 and the admin API answers with the cookie, 401 without.
- Ordered upgrades. The two planes share a database schema, so
-
The management UI knows when an operator runs it. A Deployment screen — image, replicas, policy, timeouts, workers, pool size, autoscaling, plus phase, conditions and what is actually serving — that patches the
FastllmProxyand lets the operator roll it out. It appears only under an operator: the control plane learns it is managed from an environment variable only this controller sets, so a Helm or manifest install has no such screen andGET /admin/deploymentanswers 404. The control plane reaches the API server through its own ServiceAccount and a Role naming oneresourceName, withgetandpatchand nothing else.Plus the day-1 fields a real cluster cannot do without — Service annotations (a pinned load-balancer address), scheduling, ingress, HPA,
workers/poolMaxIdle, OTLP, a ServiceMonitor, andextraArgs/extraEnvas the escape hatch — and, for the operator itself, leader election over a Lease (so it runs two replicas rather than one), Kubernetes Events, and its own/metrics,/healthzand/readyz.
Removed
- Seven of the classifier benchmarks (
potion,potion-real,potion-classes,potion-arch,potion-wide,classcheck,wrapskew). They answered "which model, which classes, which token cap" once; the answers are indocs/classifier/measurements.md, which is the artefact worth keeping.bench/minilmstays — measuring a candidate model is a question that recurs. docs/superpowers/— pre-build design notes and task lists for work that shipped, unpublished by the book and already contradicted by the code (they describe a third snapshot source that does not exist).
Fixed
POST /admin/modelssilently droppedcontext_length: the field was PATCH-only, so a caller who sent it at creation got 201 and a model with no context window. It is settable at creation now.- Every request struct in the admin API now rejects fields it does not
model. serde drops unknown fields by default, which is how the above went
unnoticed — and turning strictness on immediately found two more callers
sending
rolestoPOST /admin/principals, an endpoint that has never taken one. One of them passed["inference"]and believed it granted a role.
[0.2.0] — 2026-08-14
Two new gateways — tool servers and agents — behind the same keys, the same grants and the same accounting as models. Plus three importer bugs found by running it against a real database rather than trusting the tests.
Published as ghcr.io/azrtydxb/fastllm-proxy:v0.2.0 and
ghcr.io/azrtydxb/fastllm-operator:v0.2.0, linux/amd64 and linux/arm64.
Each architecture is built on a runner that is that architecture and the two
digests are merged into one manifest — the first attempt emulated amd64 with
QEMU and rustc segfaulted before compiling anything.
Added
- MCP gateway. One endpoint in front of every tool server, with the same
keys and the same grant machinery as models. A server is a row; a grant is
mcp:invokeonmcp/<name>and is deliberately not implied bymodel:invoke, because tools have side effects and models do not. Tools arrive namespaced<server>__<tool>so two servers can both exposesearch.GET /v1/mcp/servers,POST /v1/mcp/tools/list,POST /v1/mcp/tools/call, an MCP servers screen, and admin CRUD under/admin/mcp-servers.stdioservers are deliberately unsupported — see docs/mcp.md. - A2A gateway. One address in front of every agent:
GET /v1/agents, the agent card at/v1/agents/{name}/.well-known/…rewritten to point at this gateway so the client's next call is still authorised and attributed, andPOST /v1/agents/{name}carrying every JSON-RPC method. Protocol versions are pinned per agent rather than inferred, forwarded methods are a closed list, andagent:invokeis implied by neithermodel:invokenormcp:invoke. An Agents screen and/admin/a2a-agentsCRUD. Translation between 0.3 and 1.0 is deliberately not done — see docs/agents.md. - Interactive API reference on the docs site, rendering the same
openapi.jsonthe control plane serves at/openapi.json.
Fixed
importdropped four backend fields it had already parsed.protocol,auth_header,auth_schemeanddefault_max_tokensare declared in the config schema so a YAML file can describe an Anthropic or Azure backend, andRegistry::buildhonours them — butimportwrote three columns, so the same file produced anopenaibackend onBearerauth once it reached the database. Nothing warned: every dropped field has a valid default. Anyone who imported a native-protocol backend should re-runimport, which now converges an existing row instead of only avoiding a duplicate.- The
FileSourcepath dropped the same four, so one YAML file described a different backend depending on which code path read it. auth_schemewas two states where it needed three. Absent meansAuthorization: Bearer <key>;""means send the key with no prefix, which is what Azure'sapi-keyand Anthropic'sx-api-keyrequire. Treating absent and empty alike strippedBearerfrom every File-mode backend that did not mention the field — a request every OpenAI-compatible upstream rejects. Found by running an import against a real database and reading the row.importdropped thelimits:andbudget:blocks fromauth.keysentirely, so a key imported out of a rate-limitedFile-mode deployment arrived in the database unlimited. The key worked, which is why nobody would have looked.- A LiteLLM
anthropic/-style prefix was never stripped, soanthropic/claude-sonnet-4reached Anthropic as a model by that name. It is now stripped only when the backend speaks that protocol: to OpenRouter the same string is the model id and stripping it would ask for a model that does not exist.openrouter/joins the transport prefixes, and exactly one prefix is ever removed, soopenrouter/anthropic/claude-sonnet-4becomes the OpenRouter idanthropic/claude-sonnet-4.
[0.1.0] — 2026-08-13
First tagged release. Everything below was built before it, so this entry is a description of what 0.1.0 is rather than a diff against something earlier — grouped by capability, because there is no previous version to compare against.
Published as ghcr.io/azrtydxb/fastllm-proxy:v0.1.0 (linux/arm64).
Gateway and request path
- OpenAI-compatible gateway over any number of backends, with responses
forwarded byte-for-byte: an
openaibackend's body is never deserialised, re-encoded or buffered. - Twelve proxied
POSTendpoints:/chat/completions,/completions,/responses,/embeddings,/rerank,/score,/audio/{transcriptions,translations,speech},/images/{generations,edits}and/moderations. - Cache-affinity routing with a load escape hatch — a shared prefix returns to
the node holding its KV cache unless that node is meaningfully hotter than the
least-loaded one.
least-loadedandround-robinare selectable alternatives. - The request path performs no I/O, enforced by
tests/no_io_on_hot_path.rs. - Owned upstream connections rather than a pooled client, after the pooled one was measured as the cause of a 6× throughput difference.
- Graceful shutdown: SIGTERM stops accepting, lets in-flight generations finish
up to
--shutdown-grace(25s), and logs anything still open when it expires.
Routing
- Virtual models: ordered rules, weighted and ordered targets, and a failover chain across models, not just replicas.
- Rule conditions on principal, role, prompt and generation length, streaming, request headers, budget consumption, per-backend in-flight count, and time of day with weekday and UTC offset.
- Two-tier semantic routing — a ~115 µs static-embedding tier and an optional int8 ONNX transformer that only loads when a rule names a refined class.
POST /admin/routing/dry-runanswers which rule decided and what the chain resolved to, without dispatching anything.- A deployment-wide fallback model appended to every chain, authorised like any other candidate so it can never widen a caller's reach.
Providers and protocols
- 42 providers reachable as configuration; 40 speak the OpenAI API, 2 are translated.
- Native Anthropic (Messages) and Gemini (
generateContent) translation in both directions, including streaming, tool calls, and image and audio inputs. - Per-backend
protocol,auth_header,auth_schemeanddefault_max_tokens— reachable from both the control plane and a YAML file, which is what makes Azure OpenAI (api-keywith noBearerprefix) and native backends configurable without the database. - GCP service-account credentials minted and refreshed for Vertex AI.
Control plane, RBAC and accounting
- Control plane / data plane split behind
--role, sharing a pre-flattened snapshot;AppState::apply_snapshotis the single write path. - RBAC with real API keys: principals, roles, permissions, per-model
model:invokegrants. Keys hashed with SHA-256, passwords with Argon2id. - Upstream credentials encrypted at rest with AES-256-GCM.
- Per-principal rate limits with cross-replica reconciliation, token and spend
budgets over fixed windows, and
x-ratelimit-*response headers. - Usage accounting folded from a bounded tail buffer parsed once at end of stream; costs in integer micro-units, with prices synced from published catalogues and a provider-reported cost taking precedence.
- Append-only audit log recorded by a layer over every mutating route, with keyset pagination.
- Exact-match response cache, opt-in per model, bounded by both entries and bytes and dropped whole on any snapshot change.
Operations
- Embedded React admin UI — thirteen screens covering fleet, backends, models, routing, classes, keys, RBAC, limits, usage, audit and settings.
- Per-replica health reports over the existing proxy-token channel, surfaced by
GET /admin/fleet; kept per replica and never merged, so a partition and a dead backend stay distinguishable. - Prometheus metrics including latency histograms, cache counters and the
snapshot version, plus optional OTLP tracing behind the
otelfeature. - Reload in place: SIGHUP or a snapshot poll swaps the routing table atomically without disturbing in-flight generations.
- Runs as one binary in three shapes (
all,control,proxy); Kubernetes manifests indeploy/.
Accounting, history and the UI
- Usage recorded for every attributable request, not only for principals under a budget or a token limit — the narrower rule meant a deployment that enforced nothing recorded nothing.
- Refusals the gateway makes itself (403/429/402, and the 502 for an unreachable chain) recorded and tagged by kind, so a total backend outage no longer writes zero rows and reads as a quiet period. Unattributable refusals (401, unknown model) counted per replica per minute instead of rowed, since 401 is the one refusal a stranger can trigger at will.
GET /admin/timeseriesserves that history bucketed, with empty buckets as explicit zeros and null latency where there was nothing to measure.- 90 days of per-request rows, then hourly rollups kept indefinitely. Rollups carry no percentiles, because percentiles do not merge.
- Charts on Overview and Metrics with a click-through drill-down: five ranges, pan through history, filter by model or principal.
- Model
context_length, and a routing rule that demotes a model which cannot hold the prompt plus the requested generation. Undeclared is never treated as too small. --policy lowest-latencyfor pools whose members are not equivalent.--webhook-urlfor backend up/down and snapshot-rebuild failure, HMAC-signed./v1/modelsfiltered to what the calling key may actually invoke.
Documentation and packaging
openapi.json, served at/openapi.jsonwith Swagger UI at/docs, checked against the router in both directions bytests/openapi.rs.- A Helm chart for deployments that are not this cluster.
- Client integration guide (SDKs, five coding agents, four frameworks) and a troubleshooting page seeded from failures that actually happened.
- A Grafana dashboard and a signature-verifying webhook receiver in
examples/.
Testing
tests/protocol_fuzz.rs— mutation fuzzing over the Anthropic and Gemini translators, asserting no input panics, including arbitrary SSE chunk boundaries.tests/doc_claims.rs— countable claims in the README checked against the tables they count.web/test/— every screen mounted against wire-format fixtures, every control clicked, and the request each mutation sends asserted against the handler that receives it.- Benchmarks against LiteLLM, with the conditions and the unfavourable results
recorded alongside the favourable ones in
docs/performance.md.