Monitoring and Observability
Monitoring and observability are essential parts of operating the OnyxNet environment.
The objective is to understand whether systems are reachable, how they are performing, and where failures may be occurring.
OnyxNet uses Grafana, Prometheus, exporters, and network-specific integrations to provide visibility into infrastructure health.
Implementation Status
Section titled “Implementation Status”Status: Operational and undergoing continued development
Dashboard platform: Grafana
Metrics platform: Prometheus
Supporting components:
- Blackbox Exporter
- Node Exporter
- UniFi Poller
- Docker-based monitoring services
The monitoring environment is actively evolving. Specific integrations and dashboards may be expanded or revised over time.
Monitoring Architecture
Section titled “Monitoring Architecture”The monitoring platform separates metrics collection from visualization.
Infrastructure Targets | vMetrics Exporters | vPrometheus | vGrafana Dashboards | vOperational VisibilityPrometheus collects time-series measurements.
Grafana queries and visualizes those measurements using purpose-built dashboards.
Exporters provide the interface between monitored systems and Prometheus.
Monitoring Components
Section titled “Monitoring Components”| Component | Responsibility |
|---|---|
| Grafana | Dashboards and data visualization |
| Prometheus | Metrics collection and storage |
| Blackbox Exporter | Active network and HTTP probing |
| Node Exporter | Linux host metrics |
| UniFi Poller | UniFi network telemetry |
Different monitoring methods answer different operational questions.
A device responding to an ICMP probe does not necessarily mean every service or network path through that device is working correctly.
Network Availability
Section titled “Network Availability”Blackbox Exporter is used to perform active availability checks.
Depending on the target, these may include:
- ICMP reachability
- HTTP response checks
- Endpoint response time
- External connectivity checks
For supported probes, Prometheus exposes metrics such as:
probe_successThis indicates whether the latest probe succeeded.
A successful probe is a useful availability signal, but it is not a complete measurement of device health.
Network Latency
Section titled “Network Latency”Blackbox Exporter also exposes probe timing.
For example:
probe_duration_secondsA Grafana panel can convert seconds to milliseconds:
probe_duration_seconds * 1000This represents probe duration and should not automatically be interpreted as pure ICMP round-trip time for every probe type.
Dashboard panels must use appropriate target filters and units.
NOC-Style Dashboard
Section titled “NOC-Style Dashboard”OnyxNet includes an evolving Network Operations Center-style monitoring dashboard.
The design emphasizes quickly identifying infrastructure availability and degraded behavior.
Managed Switch Monitoring
Section titled “Managed Switch Monitoring”The environment includes three Zyxel managed switches.
Their monitoring relies on available ICMP and HTTP-based checks rather than native SNMP telemetry.
The monitoring design includes:
- Aggregate switch availability
- Individual switch reachability
- Probe duration
- Failure indicators
- Historical probe behavior
These signals help identify reachability problems but do not provide complete interface-level performance or switching telemetry.
Failure-Oriented Monitoring
Section titled “Failure-Oriented Monitoring”An important design consideration is making unhealthy states immediately visible.
For example, this expression calculates the average success value over five minutes:
avg_over_time(probe_success[5m])The result can be used to identify probes with intermittent or sustained failures.
In an actual dashboard, the query must be restricted to the relevant monitored targets.
UniFi Monitoring
Section titled “UniFi Monitoring”OnyxNet uses UniFi wireless infrastructure.
UniFi Poller provides a method of collecting telemetry from the UniFi controller and exposing metrics to Prometheus.
This allows wireless infrastructure data to participate in the broader monitoring system.
The operational objective is to combine wireless telemetry with other infrastructure metrics in one observability environment.
Sensitive controller credentials and API authentication information are not included in this public documentation.
Linux Host Monitoring
Section titled “Linux Host Monitoring”Node Exporter provides operating system metrics from supported Linux hosts.
Useful areas of monitoring include:
- CPU utilization
- Memory utilization
- Filesystem capacity
- Host availability
- Network-related metrics
Metrics should be interpreted in context.
For example, high memory utilization is not automatically a fault if most memory is being used for recoverable filesystem cache.
Monitoring Limitations
Section titled “Monitoring Limitations”Observability has practical limits.
ICMP success is not service health.
A host can respond to ping while an application is unavailable.
HTTP success is not complete application health.
A successful response may not prove that database dependencies or background jobs work.
Collection gaps affect historical accuracy.
Missing metrics may represent exporter problems, network interruptions, or target failures.
Dashboards need meaningful labels.
Metrics should identify systems consistently without embedding sensitive configuration details into public screenshots or documentation.
Alerting and Response
Section titled “Alerting and Response”Monitoring and alerting are related but distinct capabilities.
A dashboard makes information available for inspection.
Alerting evaluates conditions and directs attention toward events that may require action.
Future work should verify:
- Appropriate alert thresholds
- Avoidance of unnecessary alert noise
- Meaningful failure conditions
- Notification delivery
- Recovery notifications
- Documentation of response procedures
No specific alert-delivery guarantees are claimed without validation.
Development Roadmap
Section titled “Development Roadmap”Potential improvements include:
- Expanded infrastructure coverage
- Improved availability indicators
- Better visualization of failure states
- Consistent monitoring labels
- Additional service-level checks
- Alert testing and validation
- Improved documentation for NOC dashboards
Engineering Takeaway
Section titled “Engineering Takeaway”Observability is not just about collecting metrics.
It is about turning measurements into useful information that supports troubleshooting, operational decisions, and reliability improvements.