Appearance
Operations
Installation is covered in the Quick Start. This page is about what comes after it.
Server state
| What | Where |
|---|---|
| Service | systemctl status forgevis |
| Service journal | journalctl -u forgevis |
| Log files | /var/log/forgevis/ |
| Configuration | /etc/forgevis/config.yaml, /etc/forgevis/conf.d/ |
| Working directory | /var/lib/forgevis |
| Cluster node state | /var/lib/forgevis/data (cluster.dataDir) |
The server is ready when this answers 200:
bash
curl http://localhost:9997/api/health/readyA 503 lists the check that failed: storage, recording directory, license, cluster state. Metrics and the dashboard are in Monitoring.
Upgrading
The package installs over the one already there. A config.yaml you have changed stays yours. If the packaged config changed in the new version too, dpkg asks which one to keep and its default keeps yours; rpm puts the packaged one next to it as config.yaml.rpmnew.
Debian / Ubuntu
bash
sudo dpkg -i forgevis_*.deb
sudo systemctl restart forgevisRHEL / Rocky / AlmaLinux
bash
sudo rpm -U forgevis-*.rpm
sudo systemctl restart forgevisThe package does not restart the service itself: until it is restarted, the previous version keeps running. Recording pauses for the restart: the service stops cleanly and closes the files it was writing.
After the restart, check /api/health/ready and that the version is the new one:
bash
curl http://localhost:9997/api/health/readyCluster
Upgrade one node at a time. Move on only when the previous one is ready again (/api/health/ready answers 200) and shows as online in GET /cluster/metrics.
What happens to cameras while a node restarts:
- a node silent for more than 15 seconds counts as unavailable, and its cameras move to other nodes about 15–20 seconds after it went silent;
- if the node was the leader, the cluster elects a new one within a few seconds.
Backup
Back up the whole /etc/forgevis directory:
config.yamlandconf.d/— settings and cameras;forgevis.lic— the license file;- TLS certificates, if they are kept there.
A cluster node's state (cluster.dataDir) is not restored from a backup: a node with an outdated Raft log disrupts the cluster. The node is brought back into the cluster afresh — see below.
The recording archive is not part of this backup: its size and how long it is kept is a separate decision, see Archive retention.
Restoring
A server from scratch
- Install the package of the same version.
- Put
/etc/forgevisback from the backup. - If the archive disk survived, mount it at the same path as in
record_path. - Start the service:
sudo systemctl enable --now forgevis.
A cluster node
Recordings in the database are tied to the node's name (nodeName), and the files sit on its disks. So where possible a node goes back into the cluster as the same server, or at least under the same name — then its archive stays available.
The server is alive but the node does not work in the cluster — it does not catch up with the log, shows as offline while the process runs, or cannot come back after a failure. The node leaves the cluster with its state reset and is added back: a clean state catches up with the cluster from scratch.
- On the node itself:bashThe node takes itself out of the membership, stops its cameras and erases its Raft state. No restart is needed. If the cluster cannot accept the membership change (no leader, for example), add
curl -X POST http://<node>:9997/cluster/leave \ -H 'Content-Type: application/json' -d '{}'"force": true— then remove the node from the membership separately withPOST /cluster/remove-nodeon the leader. - Add the node back as when building a cluster:
add-learner, thenchange-membership. The node receives the state from the leader and takes cameras again.
The key and the certificate are kept. If the node does not answer the API and the state had to be removed by hand (stop the service, delete <dataDir>/node_<node_id>, start it), the key and the certificate go with it — issue a certificate again, see Certificates.
The server failed entirely. Such a node is replaced with another server: with the old one's disks or disk shelves, and under the same name.
- Remove the node from the membership:
POST /cluster/remove-nodeon the leader. Its cameras have moved by then, but the dead node's vote still counts: in a three-node cluster with one dead, one more failure leaves the cluster without a majority. - Move the archive disks or disk shelves to the new server. Give it the same
nodeNameand the samerecord_path. Until then, the node's recordings are unavailable. - Issue the node a certificate — see Certificates — and add it to the cluster as when building a cluster.
Other situations
The service does not start. The configuration is checked at startup. The reason is in journalctl -u forgevis -b: the message names the key to fix.
A cluster node is down. Its cameras move to other nodes in about 15–20 seconds. When the node is back, check /api/health/ready on it and its status in GET /cluster/metrics.
A node is cut off from the cluster. A cluster needs a majority of its voting nodes: three survive losing one, two survive losing none. A node that lost the majority keeps recording its cameras, while the rest assign the same cameras to other nodes 15–20 seconds later. Until the link is back, two nodes hold the camera: two connections to it and two recordings. If the link is gone for long, stop the service on the cut-off node:
bash
sudo systemctl stop forgevisStart it again once the link is back.
A cluster without a majority. When more than half of the voting nodes are unavailable, the cluster elects no leader and moves no cameras. The remaining nodes keep recording what was assigned to them before the failure.