Skip to content

Network and ports

Which connections an install makes, between which machines, on which ports, and which of them are encrypted. Use it to set up firewalls and to decide which networks the machines share.

At a glance

From To Port What crosses it Encrypted
Browsers, API clients Server 443 (HTTPS), 80 (redirects to HTTPS) The console, the API, the gateway (/v1) Yes, once HTTPS is on
GPU machines Server 443, or 80 while HTTPS is off The agent's connection: commands, status, downloads of the agent, runtime images and models, logs Yes, once HTTPS is on
Server Each cluster's head machine The cluster's serving port (8010 by default) Every prompt and answer, from the gateway to the models Yes, always (mutual TLS)
Server A machine running a vLLM you routed That vLLM's port plus 20000 Prompts and answers Yes, always (mutual TLS)
Server A machine running a vLLM Fadenstack found That vLLM's own port A question for the models it serves, before it is routed and when the console lists it No
Server Cluster machines Ray's metrics ports Metrics: numbers only No
Cluster machines Each other, on the cluster's network Ray's head port (6390 by default), and ports Ray picks itself Cluster control, data, calls into model copies on other machines (prompts included) Authenticated always; encrypted once the cluster's Traffic between machines is on
Cluster machines Each other, on the cluster's network GPU-to-GPU (RDMA) A model's data while one copy spans machines No: cannot be encrypted
Cluster machines Each other, on the cluster's network 8271, while a copy runs The runtime image, passed on by a machine that has it No: checked against its digest
Server The internet 443 Image pulls, model downloads, remote providers and MCP servers you connect Yes

The machines open no port for the server to control them: their agents connect out to the server and keep that connection open. To serve models, though, the server must reach the cluster's head machine on its serving port, and the machines of one cluster must reach each other directly. A machine behind NAT cannot join a cluster with machines on the other side of it.

How each of these is protected, and what to switch on: HTTPS and trust.

Users and the server

Only the server's front door faces the network: port 443 for HTTPS, and port 80, which only redirects to HTTPS once it is on. The console, the API (/api), the OpenAI-compatible gateway (/v1) and Grafana (/grafana/) are all served there.

To serve on other ports:

faden ports                            # show them, and the addresses the console is reached at
faden ports --http 8080 --https 8443   # change them; the proxy moves at once

They are FADEN_HTTP_PORT and FADEN_HTTPS_PORT in ~/fadenstack/.env. Machines and API clients then need the new port in the server's address.

Everything else on the server listens on its loopback address only and is not reachable from the network: Postgres (5432), Redis (6379), RabbitMQ (5672, 15672), ClickHouse (8123, 9000), MinIO (9090, 9091), Prometheus (9099), Loki (3100), Grafana (3001), Langfuse (3002), the gateway's own port (8001) and a few more. These ports must still be free on the server itself; faden doctor lists every port the install uses and whether something else holds it.

Warning

Do not publish those ports on the network. Loki has no authentication of its own, and Grafana's own port shows every dashboard to anyone who reaches it. The console reaches both through the front door.

The machines and the server

Each machine needs to reach the server's front door, at the address you pick in Machines → Add a machine. The panel lists the server's addresses with the interface each belongs to; pick one on the network the machine shares with the server. The server itself does not open connections to the agent.

The server does open connections to the machines for three things:

  • The models. The gateway sends every request to the head machine of the cluster that serves the model, on the cluster's serving port. Allow the server to reach that port on every head machine.
  • Metrics. The server's Prometheus collects metrics from the cluster's machines. If a firewall blocks it, the cluster page shows no metrics; serving is not affected.
  • vLLM it found. Before it routes a vLLM it found on a machine, and when the console lists those, the server asks each one which models it serves, on the vLLM's own port (see Existing vLLM servers).

Between the machines of a cluster

A cluster runs on one network that all its machines share: you confirm it when you create the cluster, and can change it under Advanced on the cluster's page. On that network the machines must reach each other freely:

  • Ray's head listens on the cluster's head port (6390 for a new cluster);
  • Ray picks further ports itself for its node managers, its data transfer and its worker processes;
  • a machine that needs the runtime image gets it from a machine of the cluster that has it, over this network;
  • GPU-to-GPU traffic (RDMA, for example RoCE between NVIDIA DGX Spark machines) runs here too.

Do not filter ports between the machines on that network. Keep the network itself to the cluster's machines: a direct cable, a switch of their own or a VLAN. GPU-to-GPU traffic cannot be encrypted, and Ray's own traffic is only encrypted once the cluster's Traffic between machines switch is on.

A machine added to a running cluster is checked first: it must reach the head on the cluster's network (see Clusters).

Ports a GPU machine must have free

Port On For
The cluster's head port (6390 by default) the head machine Ray's head
The cluster's serving port (8010 by default) the head machine the models, behind the proxy that only lets the gateway in
The serving port plus 20000 (28010 by default) the head machine, loopback only the models themselves
8271 each machine of the cluster, while it passes the runtime image on the runtime image for the other machines

Ray's dashboard listens on the head machine's loopback address only (8265). If something on a machine already holds one of these ports, the cluster cannot start and the cluster page shows Ray's error, ending in Address already in use. Free the port, then select Try again. A model published on a port another program holds is refused before it starts: "Port … is already in use on this machine by another program".

What reaches the internet

See System requirements. In short: the server pulls images and downloads models; the machines need nothing from the internet if the server holds the agent builds.