ONYXNET // ENGINEERING CASE STUDY
Infrastructure Automation
Automating Homelab Infrastructure with Ansible
Using reusable Ansible roles, YAML inventory variables, configuration management, and scheduled maintenance to operate a mixed Linux homelab.
OnyxNet contains multiple Linux systems, Raspberry Pi devices, virtualization guests, and supporting applications.
As the number of managed systems increased, manually repeating configuration and maintenance tasks became less practical.
I adopted Ansible to create a more consistent, maintainable approach to infrastructure operations.
The Engineering Problem
Manual infrastructure administration introduces several challenges.
Tasks performed successfully on one system may be repeated differently on another.
Configuration changes may go undocumented, and recovering a failed system may require reconstructing previous manual procedures.
I wanted to improve:
- Configuration consistency
- Deployment repeatability
- Operational documentation
- Maintenance efficiency
- System recoverability
- Change visibility through Git
Project Requirements
The automation environment needed to:
- Support multiple Linux platforms.
- Reuse common configuration tasks.
- Allow host-specific service selection.
- Use a centralized inventory.
- Support secure remote execution.
- Minimize unnecessary configuration changes.
- Remain understandable and maintainable.
- Accommodate Proxmox administration differences.
The objective was not to eliminate manual troubleshooting.
It was to reduce repetitive configuration work while preserving control and visibility.
Technology Selection
Ansible was selected because it supports agentless management over SSH and uses human-readable YAML for defining automation.
Git provides version history for the repository and its configuration files.
The approach can be extended gradually, starting with simple system configuration and expanding into more specialized deployment and maintenance roles.
Repository Organization
The automation repository separates responsibilities into directories.
ansible/├── inventory/├── playbooks/├── roles/├── scripts/├── docs/├── identity/└── ansible.cfgThis provides a maintainable distinction between managed systems, automation logic, supporting scripts, and documentation.
The public repository contains reusable roles for several infrastructure functions.
Reusable Roles
Examples include:
| Role | Purpose |
|---|---|
| base | Common system configuration |
| pi_base | Raspberry Pi configuration |
| docker | Container runtime deployment |
| pihole | DNS service configuration |
| dockprom | Monitoring services |
| unifi | UniFi-related deployment |
| keepalived | Availability support |
| transmission | Application deployment |
| display_kiosk | Kiosk displays |
Each role is intended to encapsulate a set of related operations.
This makes it easier to reuse automation without maintaining a separate complete playbook for every device.
Host-Specific Configuration
Not every host requires every service.
Ansible inventory variables allow different systems to use different configurations while sharing the same repository and automation structure.
For example:
enable_base: trueenable_docker: trueenable_pihole: falseenable_dockprom: falseRoles can evaluate those settings before executing their tasks.
- role: docker when: enable_docker | default(false)This approach helps avoid deploying unnecessary services to hosts that do not require them.
Idempotent Infrastructure Management
Idempotence is an important design principle in the automation environment.
A properly designed configuration task should avoid unnecessary changes when the target already meets the expected state.
Examples include:
- Ensuring a package is installed
- Creating directories with defined permissions
- Managing configuration files
- Maintaining required service states
- Mounting storage only when needed
Not every shell command is naturally idempotent.
Where commands are necessary, their behavior must be evaluated carefully.
Scheduled Automation
Some operations benefit from recurring execution rather than manual intervention.
An example from my homelab is the SimulationCraft nightly-build workflow.
Its purpose is to maintain a locally available set of recent Windows builds.
The automation includes logic to:
- Retrieve available nightly builds.
- Identify compatible Windows archives.
- Evaluate available build metadata.
- Select recent builds.
- Retain a limited number of downloads.
- Notify a Discord endpoint when appropriate.
The workflow also uses network storage, requiring mount availability to be checked during execution.
This introduced an important lesson: automation must account for the current state of dependencies before making changes.
The nightly workflow is an example from my local automation development. Its implementation may not be present in the current public repository snapshot.
Privilege Management
The environment includes systems with different administrative requirements.
Ordinary Linux hosts may require privilege escalation for system changes.
Proxmox hosts are commonly administered
through a root account and do not
necessarily provide sudo.
Applying identical privilege settings to every host can cause automation to fail.
Host-specific execution requirements must therefore be considered when designing playbooks.
Troubleshooting Automation Failures
An automation failure is not always an application failure.
Possible causes include:
- Invalid YAML or Ansible task structure
- Incorrect variable types
- File permission problems
- Network or SSH connectivity
- Missing storage mounts
- Unavailable repositories
- Incompatible operating system behavior
- Unexpected command output
Troubleshooting begins by identifying which layer failed.
For example, a scheduled task may fail before a playbook starts because the operating system cannot write to its configured log file.
That issue requires correcting file or directory permissions, not changing the application being managed.
Validation Strategy
My intended automation workflow separates several forms of verification.
Syntax and Structure
Does Ansible correctly parse the playbook?
Execution
Did the intended tasks finish successfully?
Configuration
Does the host have the expected settings?
Service Function
Is the resulting application actually usable?
Repeatability
Does running the playbook again avoid unnecessary changes?
These steps provide stronger evidence than a successful command exit code alone.
Current Results
Ansible is used as an operational configuration management platform within OnyxNet.
The environment includes a structured repository, reusable roles, host-specific settings, and maintenance automation.
The project continues to evolve.
This case study does not yet claim a measured reduction in deployment time or a formally validated recovery time for automated rebuilds.
Those measurements can be added through future testing.
Lessons Learned
Standardization Reduces Uncertainty
Consistent configuration makes systems easier to understand and maintain.
Automation Still Requires Validation
A successful playbook run does not guarantee an application is healthy.
Idempotence Matters
Repeated maintenance should not cause unnecessary configuration changes or service disruption.
Platform Differences Matter
A workflow suitable for one Linux system may require modifications for another.
Documentation Is Part of Automation
Automation code explains what actions are performed.
Operational documentation explains why they are performed and how to recover when they fail.
Future Improvements
Areas for continued development include:
- Expanding reusable roles
- Improving validation and reporting
- Strengthening configuration consistency
- Documenting service dependencies
- Improving recovery procedures
- Evaluating additional infrastructure-as-code tools
- Integrating more structured testing
Technical Documentation
Source Code
The automation repository is available at: