ONYXNET // ENGINEERING CASE STUDY
Monitoring & Observability
Building a Homelab Network Operations Dashboard
Designing a Grafana and Prometheus monitoring environment to visualize network availability, probe latency, infrastructure health, and UniFi telemetry.
Monitoring is one of the most important parts of operating OnyxNet.
As my homelab expanded to include multiple managed switches, wireless access points, servers, and virtualization environments, I wanted better visibility into infrastructure availability and network behavior.
The goal was to create a Network Operations Center-style dashboard that makes infrastructure problems easier to recognize.
The Engineering Problem
A functioning network is not necessarily a healthy network.
Individual services may fail even when the underlying host is reachable. Network devices may become unavailable, and infrastructure problems can affect multiple dependent services.
Without centralized monitoring, troubleshooting often begins only after a visible problem occurs.
I wanted to answer questions such as:
- Are all three managed switches reachable?
- Has network probe duration increased?
- Which infrastructure device is unavailable?
- Are monitored services responding correctly?
- Can wireless telemetry provide additional context during troubleshooting?
Project Requirements
The monitoring environment needed to:
- Use existing homelab hardware.
- Monitor three managed switches.
- Present useful network availability data.
- Support infrastructure and service monitoring.
- Integrate with UniFi wireless telemetry.
- Provide a clear, NOC-style dashboard.
- Remain extensible as the homelab grows.
An additional requirement was to avoid purchasing new network equipment solely to gain monitoring features.
Technology Selection
The solution uses several complementary tools.
| Component | Responsibility |
|---|---|
| Grafana | Dashboard visualization |
| Prometheus | Time-series metrics collection |
| Blackbox Exporter | Active availability probing |
| Node Exporter | Linux host telemetry |
| UniFi Poller | UniFi controller metrics |
| Docker | Monitoring service deployment |
Each component has a distinct responsibility.
Prometheus collects metrics, while Grafana provides a way to visualize and investigate them.
Exporters expose measurements from systems or perform checks against monitored targets.
Monitoring Architecture
The overall data flow is:
Monitored Infrastructure | vExporters and Probes | vPrometheus | vGrafana Dashboards | vOperational AnalysisBlackbox Exporter actively checks network and HTTP endpoints.
Node Exporter provides host-level metrics.
UniFi Poller exposes wireless infrastructure telemetry obtained from the UniFi controller.
This separation makes it possible to expand monitoring without redesigning the entire stack.
The Managed Switch Challenge
OnyxNet currently uses three Zyxel GS1200-8HPv3 managed switches.
The switches do not provide the SNMP telemetry needed for the interface-level monitoring approach I originally considered.
Rather than immediately replacing them, I adapted the monitoring approach to the capabilities of the existing hardware.
The switches can be checked through ICMP and supported HTTP-based probes.
This provides useful availability information without requiring additional purchases.
What This Approach Can Measure
- Reachability of each switch
- Whether a probe succeeds
- Probe execution duration
- Recent probe failures
- Historical availability trends
What It Cannot Fully Measure
- Per-port bandwidth utilization
- Interface error counters
- Switch CPU and memory statistics
- Complete switching-plane health
Those limitations are important.
An ICMP response confirms that an endpoint responded to a particular probe. It does not prove that every switch port, VLAN, or service is operating correctly.
Designing the NOC Dashboard
The dashboard is intended to prioritize operational information over decorative visualizations.
For switch monitoring, the design emphasizes:
- An aggregate switch availability indicator.
- Individual switch states.
- Per-target probe duration.
- Clearly visible failures.
- Historical context for troubleshooting.
For example, an aggregate indicator can communicate that all three switches are responding or identify that one has become unreachable.
The exact query must filter for the intended switch targets so unrelated probes do not affect the result.
Availability Metrics
Blackbox Exporter provides the metric:
probe_successThis commonly uses:
1for a successful probe0for a failed probe
A Grafana panel can use the metric to represent availability.
However, the result must be interpreted according to the specific probe type.
A successful ICMP probe does not establish that an HTTP application is healthy.
Failure-Oriented Monitoring
One design goal is to make failures immediately visible rather than requiring an operator to inspect every device individually.
A useful PromQL expression is:
avg_over_time(probe_success[5m])This calculates the average success value over a five-minute window.
For a continuously successful probe,
the result is generally 1.
Repeated failures reduce the value.
An inverted representation can be useful for identifying unhealthy targets.
In a production dashboard, the query must also include appropriate target labels and handle missing metrics correctly.
Measuring Probe Duration
Blackbox Exporter exposes:
probe_duration_secondsThis provides timing information for the complete probe operation.
Multiplying by 1000 expresses the result in milliseconds:
probe_duration_seconds * 1000I use timing metrics to provide additional context alongside availability data.
Probe duration is not always equivalent to pure network round-trip latency.
The distinction depends on the probe type and its individual phases.
UniFi Integration
OnyxNet also uses two UniFi AP AC-Lite wireless access points.
UniFi Poller provides telemetry from the UniFi controller and exposes Prometheus-compatible metrics.
This allows wireless information to participate in the same monitoring environment as other infrastructure.
During integration testing, metrics associated with wireless clients and access points were successfully retrieved.
The broader objective is to correlate wireless behavior with network and service availability.
Operational Lessons
Monitor Within Hardware Limitations
A useful monitoring platform does not require replacing every device that lacks advanced telemetry features.
Basic reachability measurements still provide meaningful operational information.
Understand What a Metric Proves
A green availability indicator should represent a clearly defined successful test.
It should not imply broader health guarantees than the measurement supports.
Design Dashboards for Decisions
The most useful panels are those that help answer specific operational questions.
Information should be understandable quickly enough to support troubleshooting.
Observability Is Iterative
Monitoring improves as new requirements and failure scenarios are identified.
Dashboard queries, thresholds, and visualization choices should evolve based on actual operational experience.
Current Results
The monitoring environment includes operational Grafana and Prometheus services, Blackbox probing, and UniFi Poller integration.
A NOC-style dashboard has been developed and continues to undergo refinement.
This case study does not yet include published availability statistics, measured improvements in incident response time, or a completed alerting validation report.
Those results will be added when they have been collected and verified.
Future Improvements
Planned development areas include:
- Finalizing dashboard presentation
- Expanding service-level checks
- Improving unavailable-device indicators
- Documenting alert conditions
- Testing notification delivery
- Adding sanitized dashboard screenshots
- Reviewing monitoring coverage gaps
Technical Documentation
Source Code
The public monitoring repository is available at:
The published repository represents an earlier configuration and may not reflect every component or version currently deployed.