Troubleshooting

Symptoms people actually hit, and what each one means. Most entries here are failures that happened on a real deployment rather than ones imagined for the page — where a message turned out to be misleading, that is called out, because a wrong explanation costs more than no explanation.

Requests

401 with a key you just created

The key exists in the database but the proxy has not seen it yet. Keys reach the data plane in the snapshot, which each replica polls on --config-poll (5s by default), so there is a window of a few seconds after minting.

If it persists: the key may be revoked (GET /admin/keys shows disabled), expired (expires_at — and note the UI's create form defaults to 90 days, not never), or you are sending it to the admin port instead of the data plane.

403 model_access_denied

The key is valid; the principal behind it holds no model:invoke grant for that model. Authentication and authorisation are separate, and a 403 rather than a 401 is the gateway saying so.

curl -sk -b /tmp/ck https://host:4001/admin/principals   # which roles it holds
curl -sk -b /tmp/ck https://host:4001/admin/roles        # what those roles grant

Grants are per model: model:invoke on model/<name>, or model/* for all, and the name is a frontend model's — that is what a client asks for and therefore what is granted.

A grant on a frontend model covers the whole chain it routes to. Adding a target to it extends the reach of everyone holding it, which is acceptable only because editing one needs config:write, itself enough to grant any model outright.

404 for a provider model

Provider models are inventory, not names a client may use. Ask for the frontend model in front of it — GET /v1/models lists exactly what this key can name.

404 no route for POST /v1/...

That path is not proxied. Twelve POST endpoints carrying a model are; the stateful job APIs are not, and integrations.md explains why.

429, and x-ratelimit-* headers you did not configure

Rate limits are per principal. GET /admin/limits shows who has one. Retry-After is in whole seconds and is honest — the bucket really does refill.

402 Payment Required

A budget window is exhausted. Not a rate limit: waiting does not help until the window rolls over or someone raises the cap. GET /admin/budgets.

502 upstream_unavailable

No backend in the chain could be reached. Check GET /health on the data plane for per-backend health, or the Fleet screen, which keeps replicas separate — if one replica sees a backend as down and the others do not, that is a partition rather than a dead backend, and merging them would hide it.

Requests succeed but the model returns empty content

Reasoning models put their output in reasoning_content until they finish thinking. A small max_tokens truncates the reasoning before any content appears, and finish_reason will say length. This is also the most common cause of "the model will not call tools": a tool call costs ~100 tokens of reasoning first, so a tight ceiling looks exactly like a model that ignores tools. Raise max_tokens, or disable thinking:

{ "chat_template_kwargs": { "enable_thinking": false } }

Usage, spend and charts

Usage and spend are empty

Two eras here. Before the accounting change, usage was recorded only for principals with a budget or a tokens-per-minute limit — so a deployment that enforced nothing recorded nothing. Every request is recorded now, for any authenticated caller.

If it is still empty: check that requests are reaching a backend at all (refusals are recorded separately), and that the control plane is receiving reports — fastllm_usage_reports_dropped_total on the proxy's /metrics is non-zero if the queue to the control plane is backing up.

Spend says — or unpriced

The models have no prices. A request against an unpriced model contributes nothing to a total and is counted as unpriced_requests rather than as zero cost, so a spend figure never quietly understates. Set prices with PATCH /admin/provider-models/{id} or the edit form. A self-hosted model legitimately has no price; a hosted one should have.

The chart says "a control plane older than the accounting change does not serve this"

Take that message with suspicion — it asserts a cause it has not checked. It appears whenever GET /admin/timeseries fails for any reason, including a 500. Check the endpoint directly before believing it:

curl -sk -b /tmp/ck 'https://host:4001/admin/timeseries?bucket=3600'

If that returns 500, the control-plane logs have the real reason.

A chart is empty for a window you know had traffic

Empty buckets are returned as explicit zeros, so an empty chart means no rows, not a missing series. Usage older than the retention window (90 days) has been folded into hourly rollups — the counts survive, but rolled-up buckets carry no latency, because percentiles do not merge. A latency line that stops partway back is that boundary, not a gap in traffic.

Admin plane

The browser warns about the certificate

The admin API is served with a certificate from a private CA, which no OS trust store knows. Verifying clients need the CA:

kubectl -n fastllm get secret fastllm-control-tls -o jsonpath='{.data.ca\.crt}' \
  | base64 -d > ca.crt
curl --cacert ca.crt https://host:4001/healthz

For a browser, trust that CA on the machine. Do not habitually click through certificate warnings on an admin plane.

migration N was previously applied but is missing in the resolved migrations

The binary is older than the database schema. Usually a rollback, or a manifest whose image pin drifted behind what is running. sqlx refuses rather than running against a schema it does not understand, which is the safe failure. Deploy the newer image.

A write succeeded but nothing changed

GET /admin/health reports snapshot_rebuild_failures. A write commits and then the snapshot is rebuilt; if the rebuild fails, the database and the published configuration have diverged and will stay that way until a later rebuild succeeds. This is also a webhook event, if one is configured.

One replica behaves differently from the others

Compare snapshot_version per replica on the Fleet screen. A replica on an older snapshot answers /health with ok and misbehaves only on whatever changed — most often a key it has never seen.

Disagreement right after a change is normal and is not an alert. Each proxy polls for a new snapshot every config_poll_seconds and reports its health every health_report_interval_seconds, so a replica can report the previous version for the sum of the two — 15s on the defaults — with nothing wrong.

The size of the version gap is not a measure of staleness. A version is the microsecond at which the control plane built that snapshot, and it only republishes when the content actually changed. Two consecutive versions are therefore separated by however long it happened to be between two real changes, so a replica exactly one version behind can show a gap of a second or of a minute with no difference in health. What decides it is time, by either of two signals:

  • The newest snapshot has been available longer than a poll and a report can account for (15s on the defaults, plus a small margin). Everyone has had their chance, so anyone still behind is stuck.
  • A replica has been behind continuously for longer than that. This is the one that matters on a busy gateway: Budget.tokens_used is part of the snapshot content, so traffic alone republishes it every few seconds and the newest snapshot is never old enough for the first signal to prove anything. Being behind is what persists while the version it is behind of keeps moving.

The Fleet screen says "still picking up the snapshot" until one of those holds and only raises the red banner after, reporting the snapshot's age with it. The second signal only counts time the screen has been open, so leave it open for a few polls before concluding a replica is fine.

Backends

A backend keeps being marked unhealthy

Two separate paths take a backend out of rotation, and they need different fixes, so find out which one it is before changing anything. On a build with fa2c79d or later, every ejection logs backend out of rotation with a reason; on anything older the ejection is silent and the only trace is the recovery.

  • The health probe (GET /models, unhealthy_after, default 2). A backend slow enough to exceed --health-timeout looks identical to one that is down. For long-loading engines, raise the timeout rather than the failure count. A backend ejected by real traffic (header timeouts, frozen engine counters) stays out for 30s, doubling to 5 minutes, and a passing probe does not shorten that. A model whose backends are all out is still tried, ejected ones last, when there is no further model in its chain to fail over to: a probe is a suspicion and a request is evidence, so a probe-only ejection no longer turns into 502 no healthy backend for a fleet that is serving (#29).

  • Upstream header timeouts from real traffic (consecutive_timeout_threshold, defaulting to unhealthy_after). A request that gets no response headers within upstream_timeout_seconds counts as a timeout, and that many in a row eject the backend. This is the one that bites a busy local engine: the probe answers in milliseconds while real requests queue inside vLLM long enough to time out, so the backend is ejected with nothing wrong with it. A 35B on one GPU ran a 69% 502 rate this way, serving 13k-token prompts at 37s average first byte against a default budget.

    Fix it in this order: set the backend's upstream timeout (Models page, FLOW CONTROL, or upstream_timeout_seconds on PATCH /admin/backends/{id}) above the real p99 time-to-first-byte; have the caller stream, which makes first byte arrive in seconds regardless of generation length; raise consecutive_timeout_threshold independently of unhealthy_after; and set flow control on the backend (Models page, FLOW CONTROL) so the queue forms in the proxy, bounded, instead of inside the engine.

A backend shared by several models -- a pool is a policy over an existing model's backend, not a copy of it -- is one object with one health bit. Traffic through the pool ejecting it takes the model down too, and that is correct: it is the same engine.

A backend stays out of rotation although the engine answers

If an ejected backend never logs backend healthy, back in rotation while its engine answers /v1/models in milliseconds, check that the prober is running at all. Look at the engine's own access log: each proxy should appear as a GET /metrics every --engine-scrape-interval (2s) and a GET /v1/models every --health-interval (10s). Scrapes with no probes from a proxy's node address means its prober has stopped. Nothing is logged when that happens, because nothing failed: the sweep is still waiting.

Builds before Upstream::fetch bounded only a response's headers, so one backend that sent headers and then stalled its body (a host rebooting mid-response is enough) stopped the sweep for every backend, for good. With the prober gone, a backend ejected by timeouts on every replica has nothing left that can restore it. Restarting the proxies recovers immediately, since a fresh registry starts every backend healthy. Upgrading removes the cause: every background loop now reads the whole response under one deadline.

Two nodes serving the same model behave differently

Check they are running the same engine build. A load-balanced pool whose members differ is a pool that fails intermittently and blames the router — pin by digest, not by a floating tag.

Embeddings work but their tokens are not counted

Fixed. Any non-streaming response larger than the 8 KiB tail buffer used to lose its usage, which for a 22 KB embeddings response was every one of them. If you see it on an older build, that is the cause.