Appearance
Adding nodes
A node joins the platform in two steps: it proves who it is to the platform with a password, and the platform gives it a certificate so it can prove who it is to the other nodes.
The two are separate on purpose. Managing a node is a conversation with an operator, where a password belongs. Consensus is a conversation between nodes, where an operator's password has no meaning at all.
Prepare the node
The platform authenticates as a machine user defined in the node's own configuration:
yaml
security:
auth_method: internal
users:
- user: platform
pass: "sha256:<base64 of the SHA-256 digest>"
ips: []
permissions:
- action: api
- action: metrics
- action: readEverything except the health probes sits behind this. Set the same credentials in the Control Plane when adding the server, or leave the fields empty to use the platform-wide defaults.
TIP
A wrong password is reported as a wrong password, not as an unreachable node. If a node shows as offline, the log will say which of the two it is.
Certificates
Nodes authenticate each other with certificates, and there is no switch to turn that off: anyone who can reach the consensus port of an unprotected cluster can vote, append entries and install snapshots.
Which is why a node without a certificate does not open its consensus port at all and says so:
WARN No cluster certificate yet — the Raft port stays closed until one is
issued. Add this node in the Control Plane, or point cluster.tls at
certificate files you manage yourself.This is the normal state of a fresh node, not a fault.
How one is issued
No restart. Three properties of this are load-bearing:
- The private key never leaves the node. Only a signing request does, and what comes back is a certificate, which is public by nature. The key is created once and kept, so renewal does not change the node's identity.
- The platform adds the address it uses. A node knows the address it binds; it does not know the address you reach it through — behind NAT or a published container port those differ. A certificate missing that name fails verification for exactly the caller that needs it.
- Each cluster gets its own authority. A shared root would let a node of one cluster speak consensus to another. The private key stays in the platform database, encrypted, and is never sent anywhere.
How long a certificate lasts
A node certificate is issued for 825 days; the cluster's authority for ten years. The expiry date is recorded when the certificate is issued and is visible in the fleet.
Renewal is not automatic
A certificate is issued once — when the cluster is created, or when the node joins it. Nothing re-issues on its own. A node whose certificate has expired stops passing mutual authentication and drops out of the quorum, and the log fills with handshake errors.
Renewing today means removing the node from the cluster and adding it back: issuance is tied to exactly those two operations. The private key survives leaving, and a node that returns gets a new certificate for the same key. While the node is out it is not in the cluster — plan it as maintenance, not as an instant operation.
Watch the expiry date well ahead of time. Two years later, the cause reads as a sudden loss of quorum with no configuration change behind it.
Running your own PKI
Point the node at files you manage, and the platform stays out of it:
yaml
cluster:
tls:
ca_cert: /etc/forgevis/tls/ca.pem
cert: /etc/forgevis/tls/node-1.pem
key: /etc/forgevis/tls/node-1.keyAll three together — partial configuration is refused at startup. Certificates need both serverAuth and clientAuth, because a node is a server to whoever dials it and a client to whoever it dials. The host from the target node's rpc_addr must appear in the certificate's subject alternative names.
Forming a cluster
- Add the servers to the fleet, with credentials. They appear as standalone until they belong to a cluster.
- Create a cluster from one of them. It becomes the first voter, and the platform issues its certificate before initialising it.
- Add the rest. Each gets a certificate, joins as a learner, catches up on the log, and can then be promoted to voter.
Voters elect leaders and form the quorum; learners replicate without voting. Promotion, demotion and draining a node before maintenance are all fleet operations.
Writes go to the leader
Only the leader accepts changes. A node that receives one and is not the leader answers with a redirect rather than fetching the answer itself — so nodes never need credentials for one another, and yours stay on your own request.
Removing a node
Removing the last node is not a membership change — it dissolves the cluster. Raft cannot remove its only voter, and there is no remaining membership to update, so the whole operation is telling that node to leave.
A node removed from a cluster erases its local cluster state and can be added to another one straight away, with no restart and no hand on the box.