Skip to content

Monitoring and Observability

Monitoring and observability are essential parts of operating the OnyxNet environment.

The objective is to understand whether systems are reachable, how they are performing, and where failures may be occurring.

OnyxNet uses Grafana, Prometheus, exporters, and network-specific integrations to provide visibility into infrastructure health.

Status: Operational and undergoing continued development

Dashboard platform: Grafana

Metrics platform: Prometheus

Supporting components:

  • Blackbox Exporter
  • Node Exporter
  • UniFi Poller
  • Docker-based monitoring services

The monitoring environment is actively evolving. Specific integrations and dashboards may be expanded or revised over time.

The monitoring platform separates metrics collection from visualization.

Infrastructure Targets
|
v
Metrics Exporters
|
v
Prometheus
|
v
Grafana Dashboards
|
v
Operational Visibility

Prometheus collects time-series measurements.

Grafana queries and visualizes those measurements using purpose-built dashboards.

Exporters provide the interface between monitored systems and Prometheus.

Component Responsibility
Grafana Dashboards and data visualization
Prometheus Metrics collection and storage
Blackbox Exporter Active network and HTTP probing
Node Exporter Linux host metrics
UniFi Poller UniFi network telemetry

Different monitoring methods answer different operational questions.

A device responding to an ICMP probe does not necessarily mean every service or network path through that device is working correctly.

Blackbox Exporter is used to perform active availability checks.

Depending on the target, these may include:

  • ICMP reachability
  • HTTP response checks
  • Endpoint response time
  • External connectivity checks

For supported probes, Prometheus exposes metrics such as:

probe_success

This indicates whether the latest probe succeeded.

A successful probe is a useful availability signal, but it is not a complete measurement of device health.

Blackbox Exporter also exposes probe timing.

For example:

probe_duration_seconds

A Grafana panel can convert seconds to milliseconds:

probe_duration_seconds * 1000

This represents probe duration and should not automatically be interpreted as pure ICMP round-trip time for every probe type.

Dashboard panels must use appropriate target filters and units.

OnyxNet includes an evolving Network Operations Center-style monitoring dashboard.

The design emphasizes quickly identifying infrastructure availability and degraded behavior.

The environment includes three Zyxel managed switches.

Their monitoring relies on available ICMP and HTTP-based checks rather than native SNMP telemetry.

The monitoring design includes:

  • Aggregate switch availability
  • Individual switch reachability
  • Probe duration
  • Failure indicators
  • Historical probe behavior

These signals help identify reachability problems but do not provide complete interface-level performance or switching telemetry.

An important design consideration is making unhealthy states immediately visible.

For example, this expression calculates the average success value over five minutes:

avg_over_time(probe_success[5m])

The result can be used to identify probes with intermittent or sustained failures.

In an actual dashboard, the query must be restricted to the relevant monitored targets.

OnyxNet uses UniFi wireless infrastructure.

UniFi Poller provides a method of collecting telemetry from the UniFi controller and exposing metrics to Prometheus.

This allows wireless infrastructure data to participate in the broader monitoring system.

The operational objective is to combine wireless telemetry with other infrastructure metrics in one observability environment.

Sensitive controller credentials and API authentication information are not included in this public documentation.

Node Exporter provides operating system metrics from supported Linux hosts.

Useful areas of monitoring include:

  • CPU utilization
  • Memory utilization
  • Filesystem capacity
  • Host availability
  • Network-related metrics

Metrics should be interpreted in context.

For example, high memory utilization is not automatically a fault if most memory is being used for recoverable filesystem cache.

Observability has practical limits.

ICMP success is not service health.

A host can respond to ping while an application is unavailable.

HTTP success is not complete application health.

A successful response may not prove that database dependencies or background jobs work.

Collection gaps affect historical accuracy.

Missing metrics may represent exporter problems, network interruptions, or target failures.

Dashboards need meaningful labels.

Metrics should identify systems consistently without embedding sensitive configuration details into public screenshots or documentation.

Monitoring and alerting are related but distinct capabilities.

A dashboard makes information available for inspection.

Alerting evaluates conditions and directs attention toward events that may require action.

Future work should verify:

  • Appropriate alert thresholds
  • Avoidance of unnecessary alert noise
  • Meaningful failure conditions
  • Notification delivery
  • Recovery notifications
  • Documentation of response procedures

No specific alert-delivery guarantees are claimed without validation.

Potential improvements include:

  • Expanded infrastructure coverage
  • Improved availability indicators
  • Better visualization of failure states
  • Consistent monitoring labels
  • Additional service-level checks
  • Alert testing and validation
  • Improved documentation for NOC dashboards

Observability is not just about collecting metrics.

It is about turning measurements into useful information that supports troubleshooting, operational decisions, and reliability improvements.