Troubleshooting
Symptoms people actually hit, and what each one means. Most entries here are failures that happened on a real deployment rather than ones imagined for the page — where a message turned out to be misleading, that is called out, because a wrong explanation costs more than no explanation.
Requests
401 with a key you just created
The key exists in the database but the proxy has not seen it yet. Keys reach
the data plane in the snapshot, which each replica polls on --config-poll
(5s by default), so there is a window of a few seconds after minting.
If it persists: the key may be revoked (GET /admin/keys shows disabled),
expired (expires_at — and note the UI's create form defaults to 90 days,
not never), or you are sending it to the admin port instead of the data plane.
403 model_access_denied
The key is valid; the principal behind it holds no model:invoke grant for
that model. Authentication and authorisation are separate, and a 403 rather
than a 401 is the gateway saying so.
curl -sk -b /tmp/ck https://host:4001/admin/principals # which roles it holds
curl -sk -b /tmp/ck https://host:4001/admin/roles # what those roles grant
Grants are per model: model:invoke on model/<name>, or model/* for all,
and the name is a frontend model's — that is what a client asks for and
therefore what is granted.
A grant on a frontend model covers the whole chain it routes to. Adding a
target to it extends the reach of everyone holding it, which is acceptable only
because editing one needs config:write, itself enough to grant any model
outright.
404 for a provider model
Provider models are inventory, not names a client may use. Ask for the frontend
model in front of it — GET /v1/models lists exactly what this key can name.
404 no route for POST /v1/...
That path is not proxied. Twelve POST endpoints carrying a model are; the
stateful job APIs are not, and integrations.md explains why.
429, and x-ratelimit-* headers you did not configure
Rate limits are per principal. GET /admin/limits shows who has one.
Retry-After is in whole seconds and is honest — the bucket really does refill.
402 Payment Required
A budget window is exhausted. Not a rate limit: waiting does not help until the
window rolls over or someone raises the cap. GET /admin/budgets.
502 upstream_unavailable
No backend in the chain could be reached. Check GET /health on the data plane
for per-backend health, or the Fleet screen, which keeps replicas separate —
if one replica sees a backend as down and the others do not, that is a
partition rather than a dead backend, and merging them would hide it.
Requests succeed but the model returns empty content
Reasoning models put their output in reasoning_content until they finish
thinking. A small max_tokens truncates the reasoning before any content
appears, and finish_reason will say length. This is also the most common
cause of "the model will not call tools": a tool call costs ~100 tokens of
reasoning first, so a tight ceiling looks exactly like a model that ignores
tools. Raise max_tokens, or disable thinking:
{ "chat_template_kwargs": { "enable_thinking": false } }
Usage, spend and charts
Usage and spend are empty
Two eras here. Before the accounting change, usage was recorded only for principals with a budget or a tokens-per-minute limit — so a deployment that enforced nothing recorded nothing. Every request is recorded now, for any authenticated caller.
If it is still empty: check that requests are reaching a backend at all
(refusals are recorded separately), and that the control plane is receiving
reports — fastllm_usage_reports_dropped_total on the proxy's /metrics is
non-zero if the queue to the control plane is backing up.
Spend says — or unpriced
The models have no prices. A request against an unpriced model contributes
nothing to a total and is counted as unpriced_requests rather than as zero
cost, so a spend figure never quietly understates. Set prices with
PATCH /admin/provider-models/{id} or the edit form. A self-hosted model legitimately
has no price; a hosted one should have.
The chart says "a control plane older than the accounting change does not serve this"
Take that message with suspicion — it asserts a cause it has not checked.
It appears whenever GET /admin/timeseries fails for any reason, including
a 500. Check the endpoint directly before believing it:
curl -sk -b /tmp/ck 'https://host:4001/admin/timeseries?bucket=3600'
If that returns 500, the control-plane logs have the real reason.
A chart is empty for a window you know had traffic
Empty buckets are returned as explicit zeros, so an empty chart means no rows, not a missing series. Usage older than the retention window (90 days) has been folded into hourly rollups — the counts survive, but rolled-up buckets carry no latency, because percentiles do not merge. A latency line that stops partway back is that boundary, not a gap in traffic.
Admin plane
The browser warns about the certificate
The admin API is served with a certificate from a private CA, which no OS trust store knows. Verifying clients need the CA:
kubectl -n fastllm get secret fastllm-control-tls -o jsonpath='{.data.ca\.crt}' \
| base64 -d > ca.crt
curl --cacert ca.crt https://host:4001/healthz
For a browser, trust that CA on the machine. Do not habitually click through certificate warnings on an admin plane.
migration N was previously applied but is missing in the resolved migrations
The binary is older than the database schema. Usually a rollback, or a
manifest whose image pin drifted behind what is running. sqlx refuses rather
than running against a schema it does not understand, which is the safe
failure. Deploy the newer image.
A write succeeded but nothing changed
GET /admin/health reports snapshot_rebuild_failures. A write commits and
then the snapshot is rebuilt; if the rebuild fails, the database and the
published configuration have diverged and will stay that way until a later
rebuild succeeds. This is also a webhook event, if one is configured.
One replica behaves differently from the others
Compare snapshot_version per replica on the Fleet screen. A replica on an
older snapshot answers /health with ok and misbehaves only on whatever
changed — most often a key it has never seen.
Disagreement right after a change is normal and is not an alert. Each proxy
polls for a new snapshot every config_poll_seconds and reports its health
every health_report_interval_seconds, so a replica can report the previous
version for the sum of the two — 15s on the defaults — with nothing wrong.
The size of the version gap is not a measure of staleness. A version is the microsecond at which the control plane built that snapshot, and it only republishes when the content actually changed. Two consecutive versions are therefore separated by however long it happened to be between two real changes, so a replica exactly one version behind can show a gap of a second or of a minute with no difference in health. What decides it is time, by either of two signals:
- The newest snapshot has been available longer than a poll and a report can account for (15s on the defaults, plus a small margin). Everyone has had their chance, so anyone still behind is stuck.
- A replica has been behind continuously for longer than that. This is the
one that matters on a busy gateway:
Budget.tokens_usedis part of the snapshot content, so traffic alone republishes it every few seconds and the newest snapshot is never old enough for the first signal to prove anything. Being behind is what persists while the version it is behind of keeps moving.
The Fleet screen says "still picking up the snapshot" until one of those holds and only raises the red banner after, reporting the snapshot's age with it. The second signal only counts time the screen has been open, so leave it open for a few polls before concluding a replica is fine.
Backends
A backend keeps being marked unhealthy
Two separate paths take a backend out of rotation, and they need different
fixes, so find out which one it is before changing anything. On a build with
fa2c79d or later, every ejection logs backend out of rotation with a
reason; on anything older the ejection is silent and the only trace is the
recovery.
-
The health probe (
GET /models,unhealthy_after, default 2). A backend slow enough to exceed--health-timeoutlooks identical to one that is down. For long-loading engines, raise the timeout rather than the failure count. A backend ejected by real traffic (header timeouts, frozen engine counters) stays out for 30s, doubling to 5 minutes, and a passing probe does not shorten that. A model whose backends are all out is still tried, ejected ones last, when there is no further model in its chain to fail over to: a probe is a suspicion and a request is evidence, so a probe-only ejection no longer turns into502 no healthy backendfor a fleet that is serving (#29). -
Upstream header timeouts from real traffic (
consecutive_timeout_threshold, defaulting tounhealthy_after). A request that gets no response headers withinupstream_timeout_secondscounts as a timeout, and that many in a row eject the backend. This is the one that bites a busy local engine: the probe answers in milliseconds while real requests queue inside vLLM long enough to time out, so the backend is ejected with nothing wrong with it. A 35B on one GPU ran a 69% 502 rate this way, serving 13k-token prompts at 37s average first byte against a default budget.Fix it in this order: set the backend's upstream timeout (Models page, FLOW CONTROL, or
upstream_timeout_secondsonPATCH /admin/backends/{id}) above the real p99 time-to-first-byte; have the caller stream, which makes first byte arrive in seconds regardless of generation length; raiseconsecutive_timeout_thresholdindependently ofunhealthy_after; and set flow control on the backend (Models page, FLOW CONTROL) so the queue forms in the proxy, bounded, instead of inside the engine.
A backend shared by several models -- a pool is a policy over an existing model's backend, not a copy of it -- is one object with one health bit. Traffic through the pool ejecting it takes the model down too, and that is correct: it is the same engine.
A backend stays out of rotation although the engine answers
If an ejected backend never logs backend healthy, back in rotation while its
engine answers /v1/models in milliseconds, check that the prober is running
at all. Look at the engine's own access log: each proxy should appear as a
GET /metrics every --engine-scrape-interval (2s) and a
GET /v1/models every --health-interval (10s). Scrapes with no probes from
a proxy's node address means its prober has stopped. Nothing is logged when
that happens, because nothing failed: the sweep is still waiting.
Builds before Upstream::fetch bounded only a response's headers, so one
backend that sent headers and then stalled its body (a host rebooting
mid-response is enough) stopped the sweep for every backend, for good. With
the prober gone, a backend ejected by timeouts on every replica has nothing
left that can restore it. Restarting the proxies recovers immediately, since
a fresh registry starts every backend healthy. Upgrading removes the cause:
every background loop now reads the whole response under one deadline.
Two nodes serving the same model behave differently
Check they are running the same engine build. A load-balanced pool whose members differ is a pool that fails intermittently and blames the router — pin by digest, not by a floating tag.
Embeddings work but their tokens are not counted
Fixed. Any non-streaming response larger than the 8 KiB tail buffer used to lose its usage, which for a 22 KB embeddings response was every one of them. If you see it on an older build, that is the cause.