Registering hosts that serve models
A host that serves models can tell FastLLM so, and keep telling it. When it stops, its providers stop being routed to — and eventually stop existing.
This exists because the registry drifts. Two real cases, both found by hand: a provider model pointing at a host that had been swapped to serve something else entirely, and a model with no provider at all. The first is the interesting one: the host was healthy and answering. No liveness check can catch that, because nothing was down.
What the agent does
It registers an address, on a lease, and refreshes it. That is all.
It deliberately does not send a model list. FastLLM has to reach the endpoint
anyway in order to serve traffic, so the control plane calls GET /v1/models
itself: a list pushed from the host could name models the proxies cannot dial,
and that failure would surface at request time, to a user. Enumerating from the
control plane makes discovery and reachability the same test.
It dials the control plane and is never dialled, so it works from a host behind NAT, or on a cluster FastLLM cannot reach into.
Running it
FASTLLM_CONTROL_URL=https://control.example:4001 \
FASTLLM_AGENT_TOKEN=sk-... \
python3 agent/fastllm-node-agent.py \
--advertise 192.168.10.246 \
--scan-ports 8000 8001 8890 \
--engine vllm
Standard library only, so there is nothing to install. That is deliberate: this runs on machines whose Python is whatever the vendor shipped, and a health agent that needs a virtualenv to start is one more thing to be broken at 3am.
| Flag | Why it matters |
|---|---|
--advertise | The address proxies will dial. Configured, never inferred — an agent that discovers a container on 172.17.0.2 and registers that hands the proxies an address they cannot reach. |
--scan-ports | Ports to probe on that address. Catches a bare process started by hand or by a launcher, with no container runtime present. |
--api-base | Register an endpoint outright, repeatable. Use when the address is not a port on --advertise. |
--ttl / --interval | Lease length and heartbeat. The agent refuses an interval that is not well inside the TTL, since one slow beat would then expire the lease. |
--discover-interval | How often to look for endpoints again (60s). Separate from --interval on purpose: leases on what is already known are renewed on their own clock, so a slow discovery pass never lets one lapse. |
--probe-workers | Candidates probed at once during discovery (32). A cluster exposes hundreds of ports that are not models; probed one at a time, a pass on kw took over a quarter of an hour. |
--provider-name | What this host's providers are called in FastLLM (the first part of <host>-<model>-<port>). The model and port are appended, so one host's endpoints stay distinguishable — always, not only when a second one appears, since a name that changed shape as a model was started would rename the first one behind you. Defaults to --node, and is sent on every heartbeat, so changing it renames them. |
--engine | A hint, carried as metadata. Nothing depends on it. |
--token | A principal API key. It authenticates the agent; there is no permission to grant beyond that. |
--ca-cert | PEM bundle to verify the control plane against, when its certificate comes from an internal CA. There is deliberately no way to skip verification — the token above goes over this connection, and an agent that stops checking hands it to whoever answers. Pass the CA's certificate, or a lone self-signed certificate, which is its own issuer. |
--once | Register and exit, for a cron or a smoke test. |
Running it in Kubernetes
agent/kubernetes.yaml runs the same agent with --kubernetes, where it asks
the API which addresses the cluster exposes instead of probing a list of ports.
A port probe is what you do on a host with no service registry; a cluster has
one, and guessing at ports when something can tell you is how an engine on a
port nobody thought of stays invisible.
Exposed is the whole criterion, which is why there is no label to apply and
no annotation to remember. A NodePort, a LoadBalancer and a hostPort are
reachable from outside the cluster by definition; a ClusterIP is not, and so is
never a candidate — not refused, simply not exposed, which is the answer the
operator already gave by choosing that type. Whether a candidate is a model is
then settled the same way it is for every other source: by asking it for
/v1/models. A cluster that exposes Postgres and an ingress and no engines
registers nothing.
Under --kubernetes, --scan-ports defaults to nothing. Pass it explicitly to
have both.
Host networking, and the port nobody declared
A pod on hostNetwork publishes on the node whether or not it declares a
containerPort, and one that declares none tells the API nothing about where
it listens. That is not hypothetical: this project's own engines are exactly
that — hostNetwork: true, no ports declared, serving on 8000 — and a scan
that only reads port fields finds nothing while the engine answers happily.
So for a host-network pod the agent takes, in order: a declared hostPort, then
a declared containerPort (under host networking that is a node port), then
the --port flag from the container's command, and failing all three, the
numeric port of an httpGet readiness, liveness or startup probe. The last two
read declarations the operator already wrote, which is what makes them not a
port scan — the audio servers on kw take their port from a config file, and
their readiness probe is the only place the pod spec names it. Still, they are
fallbacks. Declaring containerPort on a host-network pod is the better
fix, costs nothing, and makes the API describe the pod properly for everything
that reads it, not just this agent.
Every provider is named <cluster>-<node>-<model>-<port>: --node (the
cluster, since one agent speaks for all of it), the node the endpoint runs on,
the first model its /v1/models answers with, and the port — for example
kw-gx10-9c17-nvidia-qwen3.6-35b-a3b-nvfp4-8000 and
kw-gx10-48f4-bge-m3-8890. Each part is lowercased and anything outside
[a-z0-9.] becomes a dash. Together they stay distinct where any one part
repeats: several engines on one port, several models on one node, one model
on several nodes. A Service advertised through fastllm.io/advertise takes the
node from the pods its selector picks (all on one node, or no node part at
all); a part the agent cannot know is left out rather than guessed. The
annotation is only honoured for http:// and https:// addresses. On a bare
host --node already names the machine, so the name is
<host>-<model>-<port>. Names are sent on every heartbeat, so an upgraded
agent renames existing providers in place; routing follows model names, never
provider names.
What it needs
A ServiceAccount that can list services, pods and nodes, and nothing else —
it never writes to the cluster it runs in. Nodes because a host-network pod is
addressed by its own node's InternalIP, read from the Nodes API: a cluster runs
engines on several nodes, and one fixed address would register half of them
under the wrong host. A Deployment rather than a DaemonSet: the agent reads
cluster-wide state, so one copy answers for the cluster, where one per node
would have every node re-register the same LoadBalancer.
--advertise is optional here, and when set still means the address a
proxy will dial: for NodePorts, and as an override for a cluster whose
nodes are reached some other way than their InternalIP. It is not this pod's
address and not a ClusterIP — what is registered is a destination for someone
else's traffic, and the agent's own outbound reachability says nothing about
it.
The manifest runs the ghcr.io/azrtydxb/fastllm-node-agent image the release
build pushes on every v* tag, with the script baked in; its flags go in
args, because command would replace the image's entrypoint. It is written
for kw, where the control plane is in the same cluster and reached by its
service name, which its certificate carries. Two secrets come first:
# The control plane's CA, copied from its own TLS secret — a pod cannot
# mount a secret from another namespace.
kubectl create namespace fastllm-agent
kubectl -n fastllm get secret fastllm-control-tls -o jsonpath='{.data.ca\.crt}' \
| base64 -d > ca.crt
kubectl -n fastllm-agent create secret generic fastllm-control-ca --from-file=ca.crt
# A key for a principal of the agent's own. Any principal's key authenticates
# a registration; a dedicated one is what scopes the agent to the providers it
# registered. The proxy token is not a principal key and is refused with 401.
curl -sk -b ck -X POST https://192.168.10.129:4001/admin/principals \
-H 'content-type: application/json' -d '{"name":"node-agent-kw"}'
curl -sk -b ck -X POST https://192.168.10.129:4001/admin/keys \
-H 'content-type: application/json' \
-d '{"principal_id":"<id from above>","name":"node-agent-kw"}'
kubectl -n fastllm-agent create secret generic fastllm-agent-token \
--from-literal=token=sk-...
kubectl apply -f agent/kubernetes.yaml
Working means a registered ... leased=True line for each engine within a
minute of the pod starting, and one every --interval after that.
Running it under systemd
agent/fastllm-node-agent.service is the unit for a host outside any cluster.
This project's DGX Sparks ran it until they joined kw; it is now disabled on
both, since the in-cluster agent registers their engines and two agents
registering one endpoint fight over its name. It expects the script at /usr/local/bin/fastllm-node-agent and its
configuration in /etc/fastllm/agent.env:
FASTLLM_CONTROL_URL=https://control:4001
FASTLLM_AGENT_TOKEN=fllm_...
FASTLLM_NODE=dgx-spark
FASTLLM_ADVERTISE=192.168.10.246
sudo useradd --system --no-create-home --shell /usr/sbin/nologin fastllm-agent
sudo install -m 0755 agent/fastllm-node-agent.py /usr/local/bin/fastllm-node-agent
sudo install -d -m 0755 /etc/fastllm
sudo install -m 0644 ca.crt /etc/fastllm/ca.crt
# The token is a live credential: the agent's user, and nobody else.
sudo install -m 0640 -o root -g fastllm-agent agent.env /etc/fastllm/agent.env
sudo install -m 0644 agent/fastllm-node-agent.service /etc/systemd/system/
sudo systemctl enable --now fastllm-node-agent
Restart=always, because a host that serves models is not a host that should
stop saying so over one failed registration. The unit has no After=docker:
the agent registers whatever is serving, container or bare process, and must
come up on a host with no container runtime at all.
The name lives on the agent
A dynamic provider is named by the host registering it, not by the control
plane, which only ever sees an address. dgx-spark-8000 is a better thing to
read on a screen than 192.168.10.246:8000, and the agent is what knows which
is which.
Renaming is therefore a matter of changing --provider-name and letting the
next heartbeat carry it. That is safe: routing resolves a target by its
model's name, so a provider's name is descriptive. Nothing else refers to a
provider by name — a target names a model, and which providers serve it is the
model's business — so a rename leaves nothing dangling.
A name another provider already holds is declined — with a warning in the control plane's log — and the heartbeat still succeeds. A collision is not a reason to let a lease lapse, and an operator seeing the old name is visible and recoverable in a way that a host which quietly stopped renewing is not.
Static and cloud providers are named the other way round: a cloud provider
takes the vendor's own name from the catalogue (OpenRouter, not
openrouter.ai), and a static one is named by whoever adds it — on the
Providers screen, or with PATCH /admin/providers/{id}.
There is no RBAC on providers
A provider is an endpoint and a credential for reaching it. The credential is the provider's own — an OpenRouter key, a vLLM server's auth — defined by the provider rather than by us, and it is the whole of what a provider needs to work.
Registering one needs a token and nothing more. That is safe because registering is not an exposure: a model learned from a registered host reaches nobody until an operator points a frontend model at it, and a frontend model is where access is actually granted. A permission guarding registration would be guarding a door that opens onto nothing.
What registration deliberately cannot do is convert a provider a human typed in into one that expires — a static provider stays static, so an agent cannot take over an endpoint someone configured by hand.
That leaves a gap worth knowing about: put an agent on a host whose endpoints
were already configured by hand, and it will register them happily and change
nothing. Every line will read kind=static leased=False, and the lease, the
degradation and the model reconciliation will all be doing nothing. The
handover is an operator's explicit act:
curl -sk -b /tmp/ck -X PATCH https://control:4001/admin/providers/4 \
-H 'content-type: application/json' -d '{"kind":"dynamic"}'
and {"kind":"static"} takes it back, clearing the lease and any degradation
with it.
Any engine, in a container or not
Every engine worth naming answers GET /v1/models — vLLM, SGLang,
llama.cpp's server, TGI, Ollama, Triton's OpenAI frontend, LM Studio, mlx-lm —
as do the hosted providers. So neither the agent nor the control plane needs to
know which one it found. An unrecognised engine is registered like any other; it
simply contributes no metadata.
There is no container mode. A port probe finds a process however it was started.
What the control plane does with it
Every --provider-sweep-interval (60s by default), one GET /v1/models per
provider answers both questions that matter:
- is it reachable — the ordinary health question
- is it still serving what is registered against it — the drift question, which is the one that started all this
A provider serving more than is registered is healthy, not drifted. OpenRouter answers with hundreds of models and three of them are registered; treating the extras as drift would mark every cloud provider broken.
Degrade, then delete
A failed probe or a lapsed lease marks the provider degraded: its models go out of rotation, and nothing is deleted. Only after 30 minutes of sustained absence is a dynamic provider removed.
Two stages, never one. A 27B on a DGX Spark takes over ten minutes to load and answers nothing while it does, and a host reboot is routine. Deleting on the first failed probe would make every restart look like a decommissioning — and deleting a provider throws away its credential. Suppressing routing is reversible; deletion is not.
[!IMPORTANT] Static and cloud providers never expire. They are probed on the same schedule, but the result is advisory: it reports drift an operator would otherwise find by accident. A human put them there, and absence is not evidence the human changed their mind.
A learned model is inventory, not an exposure
A model that appears on a registered host is registered as a provider model — and reaches nobody. Nothing creates a frontend model for it.
That is deliberate. Authorisation is granted on frontend models, so a host starting an unrelated model must not hand existing principals access to it. New inventory is closed by default; exposing it is a decision someone makes, by pointing a frontend model at it.
What survives a provider being deleted
- Usage and spend.
usage_eventsrecords the model and provider name at ingest, so history outlives the row (migration 0031). - Frontend models and their targets. A target is bound by name, so deleting the provider model leaves the target in place, unresolved — and re-registering the same model on the same provider reattaches routing with no manual step (migration 0036).
- Grants, which are held on frontend models rather than on provider models.
Each of those is there because the alternative was demonstrated: the provider decomposition revoked live grants, would have failed the retention batch outright, and silently halved a frontend model's capacity — all three found by deploying and making a real request.