Appearance
Monitoring
ForgeVis provides Prometheus metrics and a ready-to-use Grafana dashboard for system monitoring.
Configuration
yaml
# Logging
logging:
directory: /var/log/forgevis # Log directory path
level: info # trace, debug, info, warn, error
console: true # Console output
# Prometheus metrics
metrics:
enabled: true
address: :9998
allowOrigin: '*'Logs:
forgevis-info-YYYY-MM-DD.log- informational logsforgevis-error-YYYY-MM-DD.log- errors and warnings- Automatic daily rotation
Metrics endpoint: http://localhost:9998/metrics
Prometheus Setup
Minimal prometheus.yml configuration:
yaml
scrape_configs:
- job_name: 'forgevis'
static_configs:
- targets: ['192.168.1.10:9998', '192.168.1.11:9998']
scrape_interval: 15sAll metrics automatically include a hostname label with the server's hostname for multi-server support.
Grafana Dashboard
Download the ForgeVis Grafana dashboard and import it into Grafana (Dashboards → New → Import → Upload JSON).
Dashboard Panels
Row 1: System Status
- Camera Status - Registered cameras / Recording now / Reconnecting / Not recording
- Camera Connections & Clients - Camera connections (main stream) / RTSP clients (main/sub) / HLS muxers (main/sub)
Row 2: Performance
- Video Frame Processing Rate - Video frame processing rate (fps)
- Broadcast Buffer Status - Internal buffer utilization (video/audio)
Row 3: Resources
- Process Memory Usage - Memory consumption (RSS)
- Average HLS Segment Size - Average HLS segment size
- Process Disk IO (Read/Write MB/s) - Per-process disk throughput (Linux only)
Row 4: Issues
- Recording Errors by Type - Recording errors by cause (connect/stream/write/internal/unclassified)
- Broadcast Lag Events & Frames Lost - Buffer lag events and dropped frames
Dashboard Variables
- Server - Select server by hostname (supports multi-select and "All")
Metrics Reference
Cameras
cameras_registered_total- Number of registered camerascameras_recording_active- Active recordingscameras_recording_errors- Cameras inside the pause before a retrycamera_streams_active_total- Active camera connections (main streams only)
RTSP
rtsp_clients_connected_total{stream="main"|"sub"}- RTSP clients by stream type
HLS
hls_muxers_active_total{stream="main"|"sub"}- Active HLS muxers
Performance
video_frames_received_total- Total video frames received (counter)broadcast_channel_capacity{channel="video"|"audio"}- Buffer capacitybroadcast_channel_len{channel="video"|"audio"}- Current buffer length
Errors
recording_errors_total{type="connect"|"stream"|"write"|"internal"|"unclassified"}- Recording errors by cause: the network or the camera, a stream gone quiet, this node's disk, a fault in the service, an error no rule recognises yet (its full text goes to the log)broadcast_lag_events_total{consumer="recorder"|"rtsp"|"hls"}- Lag events by consumerbroadcast_lag_frames_total- Frames dropped due to lag
System
process_resident_memory_bytes- Process memory usage (RSS)process_start_time_seconds- Process start time (Unix timestamp)process_cpu_seconds_total- Total CPU time used by process (seconds)recording_segments_written_total- Total segments writtenrecord_fs_total_bytes{path_root="..."}- Recording filesystem total bytes by path rootrecord_fs_free_bytes{path_root="..."}- Recording filesystem free bytes by path root
Cluster
cluster_nodes_total- Total nodes in cluster membershipcluster_nodes_online_total- Online nodes by heartbeat // Linux-only process I/O metricsprocess_io_read_bytes_total- Bytes read by the ForgeVis process (from storage)process_io_write_bytes_total- Bytes written by the ForgeVis process (to storage)process_io_cancelled_write_bytes_total- Bytes of cancelled writes (e.g. truncation, page cache invalidation)
All metrics automatically include {hostname="..."} label for server filtering.
Multi-Server Monitoring
When running multiple ForgeVis instances:
- Each server automatically reports its hostname via
hostname::get() - Prometheus collects metrics from all servers with their unique hostnames
- Grafana dashboard provides a Server dropdown to filter or view all servers
- Select "All" to see aggregated metrics across all servers
- Select specific servers to troubleshoot individual instances
Memory Usage Example:
- Server shows
8.75 MBviaps aux→ Metric shows9158656 bytes(~8.7 MB) - Dashboard displays in KB/MB using
decbytesunit for human-readable format