◆ ForgeVis
Skip to content

Monitoring ​

ForgeVis provides Prometheus metrics and a ready-to-use Grafana dashboard for system monitoring.

Configuration ​

yaml
# Logging
logging:
  directory: /var/log/forgevis  # Log directory path
  level: info                       # trace, debug, info, warn, error
  console: true                     # Console output

# Prometheus metrics
metrics:
  enabled: true
  address: :9998
  allowOrigin: '*'

Logs:

  • forgevis-info-YYYY-MM-DD.log - informational logs
  • forgevis-error-YYYY-MM-DD.log - errors and warnings
  • Automatic daily rotation

Metrics endpoint: http://localhost:9998/metrics

Prometheus Setup ​

Minimal prometheus.yml configuration:

yaml
scrape_configs:
  - job_name: 'forgevis'
    static_configs:
      - targets: ['192.168.1.10:9998', '192.168.1.11:9998']
    scrape_interval: 15s

All metrics automatically include a hostname label with the server's hostname for multi-server support.

Grafana Dashboard ​

Download the ForgeVis Grafana dashboard and import it into Grafana (Dashboards → New → Import → Upload JSON).

Dashboard Panels ​

Row 1: System Status

  • Camera Status - Registered cameras / Recording now / Reconnecting / Not recording
  • Camera Connections & Clients - Camera connections (main stream) / RTSP clients (main/sub) / HLS muxers (main/sub)

Row 2: Performance

  • Video Frame Processing Rate - Video frame processing rate (fps)
  • Broadcast Buffer Status - Internal buffer utilization (video/audio)

Row 3: Resources

  • Process Memory Usage - Memory consumption (RSS)
  • Average HLS Segment Size - Average HLS segment size
  • Process Disk IO (Read/Write MB/s) - Per-process disk throughput (Linux only)

Row 4: Issues

  • Recording Errors by Type - Recording errors by cause (connect/stream/write/internal/unclassified)
  • Broadcast Lag Events & Frames Lost - Buffer lag events and dropped frames

Dashboard Variables ​

  • Server - Select server by hostname (supports multi-select and "All")

Metrics Reference ​

Cameras ​

  • cameras_registered_total - Number of registered cameras
  • cameras_recording_active - Active recordings
  • cameras_recording_errors - Cameras inside the pause before a retry
  • camera_streams_active_total - Active camera connections (main streams only)

RTSP ​

  • rtsp_clients_connected_total{stream="main"|"sub"} - RTSP clients by stream type

HLS ​

  • hls_muxers_active_total{stream="main"|"sub"} - Active HLS muxers

Performance ​

  • video_frames_received_total - Total video frames received (counter)
  • broadcast_channel_capacity{channel="video"|"audio"} - Buffer capacity
  • broadcast_channel_len{channel="video"|"audio"} - Current buffer length

Errors ​

  • recording_errors_total{type="connect"|"stream"|"write"|"internal"|"unclassified"} - Recording errors by cause: the network or the camera, a stream gone quiet, this node's disk, a fault in the service, an error no rule recognises yet (its full text goes to the log)
  • broadcast_lag_events_total{consumer="recorder"|"rtsp"|"hls"} - Lag events by consumer
  • broadcast_lag_frames_total - Frames dropped due to lag

System ​

  • process_resident_memory_bytes - Process memory usage (RSS)
  • process_start_time_seconds - Process start time (Unix timestamp)
  • process_cpu_seconds_total - Total CPU time used by process (seconds)
  • recording_segments_written_total - Total segments written
  • record_fs_total_bytes{path_root="..."} - Recording filesystem total bytes by path root
  • record_fs_free_bytes{path_root="..."} - Recording filesystem free bytes by path root

Cluster ​

  • cluster_nodes_total - Total nodes in cluster membership
  • cluster_nodes_online_total - Online nodes by heartbeat // Linux-only process I/O metrics
  • process_io_read_bytes_total - Bytes read by the ForgeVis process (from storage)
  • process_io_write_bytes_total - Bytes written by the ForgeVis process (to storage)
  • process_io_cancelled_write_bytes_total - Bytes of cancelled writes (e.g. truncation, page cache invalidation)

All metrics automatically include {hostname="..."} label for server filtering.

Multi-Server Monitoring ​

When running multiple ForgeVis instances:

  1. Each server automatically reports its hostname via hostname::get()
  2. Prometheus collects metrics from all servers with their unique hostnames
  3. Grafana dashboard provides a Server dropdown to filter or view all servers
  4. Select "All" to see aggregated metrics across all servers
  5. Select specific servers to troubleshoot individual instances

Memory Usage Example:

  • Server shows 8.75 MB via ps aux → Metric shows 9158656 bytes (~8.7 MB)
  • Dashboard displays in KB/MB using decbytes unit for human-readable format

Proprietary software.