Operate¶
This section is for whoever runs Fadenstack: the server, the GPU machines and the clusters they form. It covers installing, upgrading, securing and backing up the server, connecting machines, and keeping it all running.
What you run¶
The server. One Linux machine that runs Fadenstack as a set of Docker containers: the console, the API, the
OpenAI-compatible gateway (/v1), the database, and the logs and metrics stores. It needs no GPU. You install
and upgrade it with the faden command-line tool.
GPU machines. Linux machines with NVIDIA GPUs that run the models. Each one runs a small program, the node
agent (faden-agent), that connects out to the server and carries out what the console asks of it. You add a
machine by running one command on it and approving it in the console.
Clusters. A cluster groups one or more machines so they can serve models together. The machines run the models inside a runtime image: a container with Ray, which spreads the work over the machines, and vLLM, the engine that runs a model. You create and manage clusters in the console; the models themselves are deployed by an administrator (see Deployments).
A cluster runs on its machines, not on the server. Its models keep running while the server restarts or is upgraded; only requests through the gateway wait until the gateway is back.
The order of work¶
- Check the System requirements for the server and the machines.
- Install the server and sign in.
- Decide on the certificate: HTTPS and trust. A new install serves HTTPS from the start.
- Open the ports between users, the server and the machines.
- Add the GPU machines.
- Create a cluster from them.
- Hand over to the administrator, who deploys models and adds people (see Administer).
- Keep backups, and upgrade when a release comes out.
When something goes wrong¶
- Failures and logs: what recovers by itself, what you see while it does, and where the logs are.
- Troubleshooting: common problems, their cause, and what to do.
Also here¶
- Existing vLLM servers: put a vLLM you already run behind the gateway.
- Taking over clusters: a rebuilt server takes over the clusters its machines still run.
- Containers: the Docker containers on the server, as the console shows them.
- Third-party components: what the server and the runtime images are built from.