System requirements¶
What the server and the GPU machines need before you install anything. The server is sized by the services it runs and the people using them; a GPU machine is sized by the models it runs.
The server¶
| Minimum | Recommended | |
|---|---|---|
| Operating system | Linux, x86_64 | |
| CPU cores | 4 | 8 |
| Memory, core services | 6 GB | 8 GB |
| Memory, with the PII, MCP and skills modules | 8 GB | 12 GB |
| Disk | 60 GB | 100 GB, plus the size of the models you intend to host |
| GPU | none | none |
- Docker Engine 24 or later with Compose v2. The user who runs
fadenmust be able to use Docker (be in thedockergroup). - Python 3.12 or later, with
pipxoruv, to install thefadencommand-line tool. - Ports 80 and 443 free on the server. Other ports can be chosen; see Network and ports.
- Internet access while installing and upgrading, to pull the service images and the runtime images. See What needs the internet.
Add memory for anything else you run on the server. The server needs no GPU: the models run on the machines you add. A server that has an NVIDIA GPU can also be added as a machine.
faden doctor checks a host before you install: memory, free disk, Docker, Compose, and whether the ports the
install publishes are free.
What grows with what¶
| When you add | What grows on the server |
|---|---|
| People chatting at the same time | CPU of the gateway and the console API first, then database connections |
| The PII module | About 2 GB of memory (it loads a language model), more CPU per chat turn |
| MCP tools, skills | A few hundred MB of memory |
| Models | Disk: every model is downloaded to the server and copied to the machines from there |
| GPU machines | One runtime image per kind of machine on disk; nothing measurable in CPU or memory |
| Tracing | Disk for the trace store |
faden deploy and faden upgrade size the services for the host they run on: how many worker processes each
service gets, and how many database connections Postgres accepts. To size them again after changing the
hardware, run faden tune and then faden up.
Disk¶
| What | Size | Needed when |
|---|---|---|
| Service images | about 11 GB | always |
| Runtime image, x86_64 machines | about 15 GB | you add x86_64 GPU machines |
| Runtime image, NVIDIA DGX Spark machines | about 12 GB | you add DGX Spark machines |
| Model store | the models you host | always |
| Database | grows with chat history | always |
| Trace store | grows with traces | tracing is on |
The server keeps a runtime image for each kind of machine you have, because the machines download it from the
server, not from the internet. If you only have x86_64 machines, install with
faden deploy --runtime-images x86_64 to skip the other one.
Backups land on the same disk by default; leave room for them, or send them elsewhere (see Backups).
GPU machines¶
| Operating system | Linux: x86_64, or aarch64 on NVIDIA DGX Spark only |
| GPU | NVIDIA only, with a working driver: nvidia-smi must print your card |
| Containers | Docker or Podman, running, and usable by the user the agent runs as, with access to the GPU (for Docker, the NVIDIA Container Toolkit) |
| Service manager | systemd: the agent runs as a system service, or as a user service when installed without sudo |
| Disk | about 15 GB for the runtime image, plus room for the models it runs; a runtime switch needs room for a second image while it runs |
| Network | Reach the server's HTTPS port. The server must reach the machine's model port. Machines in one cluster must reach each other. See Network and ports |
Two kinds of machine are supported, each with its own runtime image:
- x86_64 machines with NVIDIA GPUs, from the Turing generation to Blackwell. Their runtime image is not certified yet: the cluster's Runtime card says Not certified.
- NVIDIA DGX Spark (GB10, aarch64).
Other machines are not supported. One without an NVIDIA GPU, or with another processor architecture, joins but cannot be put in a cluster: the cluster wizard says that no certified runtime image runs on it. Another aarch64 machine with an NVIDIA GPU is offered the DGX Spark image, which is not made for it.
Note
A model runs on one kind of GPU at a time. Mixed machines can share a cluster, but each copy of a model lands on machines of one kind. Machines that should run one model together need the same GPUs.
GPU memory decides which models fit. The model marketplace shows, for each model and cluster, whether it fits on a share of one accelerator, on one, across several, or not at all (see Models).
What needs the internet¶
| Who | Reaches | When |
|---|---|---|
| Server | Container registries (ghcr.io, Docker Hub and the others the images come from) |
installing and upgrading |
| Server | Hugging Face | downloading models |
| Server | Remote model providers, remote MCP servers | if the administrator connects any |
| Server | GitHub | to check its own copy of the agent builds, if it holds one |
| GPU machines | GitHub (the project's releases) | installing or updating the agent, unless the server holds the agent builds |
| GPU machines | Hugging Face | only on a cluster set to let its machines download models themselves |
Everything else a machine needs, the runtime image and the models, comes from the server. A machine without internet access works as long as the server holds the agent builds; see Machines without internet access.