Registering hosts that serve models

A host that serves models can tell FastLLM so, and keep telling it. When it stops, its providers stop being routed to — and eventually stop existing.

This exists because the registry drifts. Two real cases, both found by hand: a provider model pointing at a host that had been swapped to serve something else entirely, and a model with no provider at all. The first is the interesting one: the host was healthy and answering. No liveness check can catch that, because nothing was down.

What the agent does

It registers an address, on a lease, and refreshes it. That is all.

It deliberately does not send a model list. FastLLM has to reach the endpoint anyway in order to serve traffic, so the control plane calls GET /v1/models itself: a list pushed from the host could name models the proxies cannot dial, and that failure would surface at request time, to a user. Enumerating from the control plane makes discovery and reachability the same test.

It dials the control plane and is never dialled, so it works from a host behind NAT, or on a cluster FastLLM cannot reach into.

Running it

FASTLLM_CONTROL_URL=https://control.example:4001 \
FASTLLM_AGENT_TOKEN=sk-... \
python3 agent/fastllm-node-agent.py \
  --advertise 192.168.10.246 \
  --scan-ports 8000 8001 8890 \
  --engine vllm

Standard library only, so there is nothing to install. That is deliberate: this runs on machines whose Python is whatever the vendor shipped, and a health agent that needs a virtualenv to start is one more thing to be broken at 3am.

FlagWhy it matters
--advertiseThe address proxies will dial. Configured, never inferred — an agent that discovers a container on 172.17.0.2 and registers that hands the proxies an address they cannot reach.
--scan-portsPorts to probe on that address. Catches a bare process started by hand or by a launcher, with no container runtime present.
--api-baseRegister an endpoint outright, repeatable. Use when the address is not a port on --advertise.
--ttl / --intervalLease length and heartbeat. The agent refuses an interval that is not well inside the TTL, since one slow beat would then expire the lease.
--discover-intervalHow often to look for endpoints again (60s). Separate from --interval on purpose: leases on what is already known are renewed on their own clock, so a slow discovery pass never lets one lapse.
--probe-workersCandidates probed at once during discovery (32). A cluster exposes hundreds of ports that are not models; probed one at a time, a pass on kw took over a quarter of an hour.
--provider-nameWhat this host's providers are called in FastLLM (the first part of <host>-<model>-<port>). The model and port are appended, so one host's endpoints stay distinguishable — always, not only when a second one appears, since a name that changed shape as a model was started would rename the first one behind you. Defaults to --node, and is sent on every heartbeat, so changing it renames them.
--engineA hint, carried as metadata. Nothing depends on it.
--tokenA principal API key. It authenticates the agent; there is no permission to grant beyond that.
--ca-certPEM bundle to verify the control plane against, when its certificate comes from an internal CA. There is deliberately no way to skip verification — the token above goes over this connection, and an agent that stops checking hands it to whoever answers. Pass the CA's certificate, or a lone self-signed certificate, which is its own issuer.
--onceRegister and exit, for a cron or a smoke test.

Running it in Kubernetes

agent/kubernetes.yaml runs the same agent with --kubernetes, where it asks the API which addresses the cluster exposes instead of probing a list of ports. A port probe is what you do on a host with no service registry; a cluster has one, and guessing at ports when something can tell you is how an engine on a port nobody thought of stays invisible.

Exposed is the whole criterion, which is why there is no label to apply and no annotation to remember. A NodePort, a LoadBalancer and a hostPort are reachable from outside the cluster by definition; a ClusterIP is not, and so is never a candidate — not refused, simply not exposed, which is the answer the operator already gave by choosing that type. Whether a candidate is a model is then settled the same way it is for every other source: by asking it for /v1/models. A cluster that exposes Postgres and an ingress and no engines registers nothing.

Under --kubernetes, --scan-ports defaults to nothing. Pass it explicitly to have both.

Host networking, and the port nobody declared

A pod on hostNetwork publishes on the node whether or not it declares a containerPort, and one that declares none tells the API nothing about where it listens. That is not hypothetical: this project's own engines are exactly that — hostNetwork: true, no ports declared, serving on 8000 — and a scan that only reads port fields finds nothing while the engine answers happily.

So for a host-network pod the agent takes, in order: a declared hostPort, then a declared containerPort (under host networking that is a node port), then the --port flag from the container's command, and failing all three, the numeric port of an httpGet readiness, liveness or startup probe. The last two read declarations the operator already wrote, which is what makes them not a port scan — the audio servers on kw take their port from a config file, and their readiness probe is the only place the pod spec names it. Still, they are fallbacks. Declaring containerPort on a host-network pod is the better fix, costs nothing, and makes the API describe the pod properly for everything that reads it, not just this agent.

Every provider is named <cluster>-<node>-<model>-<port>: --node (the cluster, since one agent speaks for all of it), the node the endpoint runs on, the first model its /v1/models answers with, and the port — for example kw-gx10-9c17-nvidia-qwen3.6-35b-a3b-nvfp4-8000 and kw-gx10-48f4-bge-m3-8890. Each part is lowercased and anything outside [a-z0-9.] becomes a dash. Together they stay distinct where any one part repeats: several engines on one port, several models on one node, one model on several nodes. A Service advertised through fastllm.io/advertise takes the node from the pods its selector picks (all on one node, or no node part at all); a part the agent cannot know is left out rather than guessed. The annotation is only honoured for http:// and https:// addresses. On a bare host --node already names the machine, so the name is <host>-<model>-<port>. Names are sent on every heartbeat, so an upgraded agent renames existing providers in place; routing follows model names, never provider names.

What it needs

A ServiceAccount that can list services, pods and nodes, and nothing else — it never writes to the cluster it runs in. Nodes because a host-network pod is addressed by its own node's InternalIP, read from the Nodes API: a cluster runs engines on several nodes, and one fixed address would register half of them under the wrong host. A Deployment rather than a DaemonSet: the agent reads cluster-wide state, so one copy answers for the cluster, where one per node would have every node re-register the same LoadBalancer.

--advertise is optional here, and when set still means the address a proxy will dial: for NodePorts, and as an override for a cluster whose nodes are reached some other way than their InternalIP. It is not this pod's address and not a ClusterIP — what is registered is a destination for someone else's traffic, and the agent's own outbound reachability says nothing about it.

The manifest runs the ghcr.io/azrtydxb/fastllm-node-agent image the release build pushes on every v* tag, with the script baked in; its flags go in args, because command would replace the image's entrypoint. It is written for kw, where the control plane is in the same cluster and reached by its service name, which its certificate carries. Two secrets come first:

# The control plane's CA, copied from its own TLS secret — a pod cannot
# mount a secret from another namespace.
kubectl create namespace fastllm-agent
kubectl -n fastllm get secret fastllm-control-tls -o jsonpath='{.data.ca\.crt}' \
  | base64 -d > ca.crt
kubectl -n fastllm-agent create secret generic fastllm-control-ca --from-file=ca.crt

# A key for a principal of the agent's own. Any principal's key authenticates
# a registration; a dedicated one is what scopes the agent to the providers it
# registered. The proxy token is not a principal key and is refused with 401.
curl -sk -b ck -X POST https://192.168.10.129:4001/admin/principals \
  -H 'content-type: application/json' -d '{"name":"node-agent-kw"}'
curl -sk -b ck -X POST https://192.168.10.129:4001/admin/keys \
  -H 'content-type: application/json' \
  -d '{"principal_id":"<id from above>","name":"node-agent-kw"}'
kubectl -n fastllm-agent create secret generic fastllm-agent-token \
  --from-literal=token=sk-...

kubectl apply -f agent/kubernetes.yaml

Working means a registered ... leased=True line for each engine within a minute of the pod starting, and one every --interval after that.

Running it under systemd

agent/fastllm-node-agent.service is the unit for a host outside any cluster. This project's DGX Sparks ran it until they joined kw; it is now disabled on both, since the in-cluster agent registers their engines and two agents registering one endpoint fight over its name. It expects the script at /usr/local/bin/fastllm-node-agent and its configuration in /etc/fastllm/agent.env:

FASTLLM_CONTROL_URL=https://control:4001
FASTLLM_AGENT_TOKEN=fllm_...
FASTLLM_NODE=dgx-spark
FASTLLM_ADVERTISE=192.168.10.246
sudo useradd --system --no-create-home --shell /usr/sbin/nologin fastllm-agent
sudo install -m 0755 agent/fastllm-node-agent.py /usr/local/bin/fastllm-node-agent
sudo install -d -m 0755 /etc/fastllm
sudo install -m 0644 ca.crt /etc/fastllm/ca.crt
# The token is a live credential: the agent's user, and nobody else.
sudo install -m 0640 -o root -g fastllm-agent agent.env /etc/fastllm/agent.env
sudo install -m 0644 agent/fastllm-node-agent.service /etc/systemd/system/
sudo systemctl enable --now fastllm-node-agent

Restart=always, because a host that serves models is not a host that should stop saying so over one failed registration. The unit has no After=docker: the agent registers whatever is serving, container or bare process, and must come up on a host with no container runtime at all.

The name lives on the agent

A dynamic provider is named by the host registering it, not by the control plane, which only ever sees an address. dgx-spark-8000 is a better thing to read on a screen than 192.168.10.246:8000, and the agent is what knows which is which.

Renaming is therefore a matter of changing --provider-name and letting the next heartbeat carry it. That is safe: routing resolves a target by its model's name, so a provider's name is descriptive. Nothing else refers to a provider by name — a target names a model, and which providers serve it is the model's business — so a rename leaves nothing dangling.

A name another provider already holds is declined — with a warning in the control plane's log — and the heartbeat still succeeds. A collision is not a reason to let a lease lapse, and an operator seeing the old name is visible and recoverable in a way that a host which quietly stopped renewing is not.

Static and cloud providers are named the other way round: a cloud provider takes the vendor's own name from the catalogue (OpenRouter, not openrouter.ai), and a static one is named by whoever adds it — on the Providers screen, or with PATCH /admin/providers/{id}.

There is no RBAC on providers

A provider is an endpoint and a credential for reaching it. The credential is the provider's own — an OpenRouter key, a vLLM server's auth — defined by the provider rather than by us, and it is the whole of what a provider needs to work.

Registering one needs a token and nothing more. That is safe because registering is not an exposure: a model learned from a registered host reaches nobody until an operator points a frontend model at it, and a frontend model is where access is actually granted. A permission guarding registration would be guarding a door that opens onto nothing.

What registration deliberately cannot do is convert a provider a human typed in into one that expires — a static provider stays static, so an agent cannot take over an endpoint someone configured by hand.

That leaves a gap worth knowing about: put an agent on a host whose endpoints were already configured by hand, and it will register them happily and change nothing. Every line will read kind=static leased=False, and the lease, the degradation and the model reconciliation will all be doing nothing. The handover is an operator's explicit act:

curl -sk -b /tmp/ck -X PATCH https://control:4001/admin/providers/4 \
  -H 'content-type: application/json' -d '{"kind":"dynamic"}'

and {"kind":"static"} takes it back, clearing the lease and any degradation with it.

Any engine, in a container or not

Every engine worth naming answers GET /v1/models — vLLM, SGLang, llama.cpp's server, TGI, Ollama, Triton's OpenAI frontend, LM Studio, mlx-lm — as do the hosted providers. So neither the agent nor the control plane needs to know which one it found. An unrecognised engine is registered like any other; it simply contributes no metadata.

There is no container mode. A port probe finds a process however it was started.

What the control plane does with it

Every --provider-sweep-interval (60s by default), one GET /v1/models per provider answers both questions that matter:

  • is it reachable — the ordinary health question
  • is it still serving what is registered against it — the drift question, which is the one that started all this

A provider serving more than is registered is healthy, not drifted. OpenRouter answers with hundreds of models and three of them are registered; treating the extras as drift would mark every cloud provider broken.

Degrade, then delete

A failed probe or a lapsed lease marks the provider degraded: its models go out of rotation, and nothing is deleted. Only after 30 minutes of sustained absence is a dynamic provider removed.

Two stages, never one. A 27B on a DGX Spark takes over ten minutes to load and answers nothing while it does, and a host reboot is routine. Deleting on the first failed probe would make every restart look like a decommissioning — and deleting a provider throws away its credential. Suppressing routing is reversible; deletion is not.

[!IMPORTANT] Static and cloud providers never expire. They are probed on the same schedule, but the result is advisory: it reports drift an operator would otherwise find by accident. A human put them there, and absence is not evidence the human changed their mind.

A learned model is inventory, not an exposure

A model that appears on a registered host is registered as a provider model — and reaches nobody. Nothing creates a frontend model for it.

That is deliberate. Authorisation is granted on frontend models, so a host starting an unrelated model must not hand existing principals access to it. New inventory is closed by default; exposing it is a decision someone makes, by pointing a frontend model at it.

What survives a provider being deleted

  • Usage and spend. usage_events records the model and provider name at ingest, so history outlives the row (migration 0031).
  • Frontend models and their targets. A target is bound by name, so deleting the provider model leaves the target in place, unresolved — and re-registering the same model on the same provider reattaches routing with no manual step (migration 0036).
  • Grants, which are held on frontend models rather than on provider models.

Each of those is there because the alternative was demonstrated: the provider decomposition revoked live grants, would have failed the retention batch outright, and silently halved a frontend model's capacity — all three found by deploying and making a real request.