Skip to content

Ansible Infrastructure Automation

Ansible is a key part of the OnyxNet infrastructure management strategy.

It provides a centralized, version-controlled approach to Linux configuration, service deployment, and repeatable operational tasks.

Rather than relying exclusively on manually executed commands, the goal is to express infrastructure configuration through reusable automation.

Platform: Ansible

Primary purpose: Configuration management and automation

Repository: OnyxNet Ansible

Status: In use and under continued development

The published repository provides a reference for the automation structure. Local configurations and newer playbooks may differ from the published revision.

OnyxNet includes Linux systems, Raspberry Pi devices, virtualization guests, and supporting infrastructure services.

Managing these systems manually introduces challenges:

  • Configuration drift
  • Inconsistent deployments
  • Repetitive maintenance
  • Difficulty reproducing a working configuration
  • Incomplete operational documentation

Ansible helps address these challenges by expressing configuration requirements in readable YAML files.

The automation repository is organized by responsibility.

ansible/
├── ansible.cfg
├── inventory/
├── playbooks/
├── roles/
├── scripts/
├── docs/
└── identity/
Component Purpose
Inventory Defines managed hosts and groups
Playbooks Orchestrates configuration tasks
Roles Groups reusable automation
Scripts Provides supporting utilities
Documentation Records architecture and operational procedures
Identity Supports SSH identity management

The public portfolio does not include private credentials or complete production inventory data.

The published repository includes roles covering several infrastructure functions.

Role Responsibility
base Common system configuration
pi_base Raspberry Pi-specific setup
docker Docker installation and configuration
dockprom Monitoring-stack deployment
pihole DNS filtering services
keepalived High-availability support
unifi UniFi-related services
transmission Download service deployment
display_kiosk Kiosk display configuration
display_oled OLED display integration

These are repository capabilities, not a claim that every role is currently enabled on every host.

OnyxNet uses inventory variables to control which roles apply to individual systems.

For example, an inventory host may define:

enable_base: true
enable_docker: true
enable_dockprom: false
enable_pihole: false

A playbook can use these variables to conditionally include roles:

- role: docker
when: enable_docker | default(false)
- role: dockprom
when: enable_dockprom | default(false)

This approach allows a shared automation repository to support systems with different responsibilities.

Infrastructure automation follows a structured process.

Identify Requirement
|
v
Review Existing Configuration
|
v
Update Role or Playbook
|
v
Validate Changes
|
v
Deploy to Selected Hosts
|
v
Verify Intended Results
|
v
Document Outcome

A successful playbook execution is not, by itself, proof that the application or service works correctly.

Post-deployment validation is still necessary.

An important Ansible design objective is idempotence: running the same playbook repeatedly should not introduce unnecessary changes when the target system already matches its desired configuration.

Typical examples include:

  • Creating a directory only when required
  • Managing file ownership and permissions
  • Installing packages declaratively
  • Managing configuration files
  • Restarting services only when necessary

Commands and scripts that always report changes require additional attention.

Idempotence should be tested rather than assumed.

Different managed operating systems have different privilege requirements.

Many Linux hosts use privilege escalation through become.

Proxmox hosts are generally administered directly as root in this environment and do not use sudo in the same way as ordinary Linux clients.

Playbooks targeting Proxmox must account for that difference rather than blindly applying the same privilege escalation settings to every host.

Automation credentials and sensitive configuration must be protected.

Operational practices should include:

  • SSH key-based authentication
  • Appropriate private-key permissions
  • Separation of public code and sensitive variables
  • Restricted access to automation credentials
  • Reviewing configuration before publication
  • Avoiding credentials in logs and Git history

Public documentation describes the architecture without exposing production secrets.

A reliable automation workflow should distinguish between several forms of validation.

Syntax validation

Does the playbook parse correctly?

Execution validation

Did Ansible complete the intended tasks?

Configuration validation

Does the target match the desired configuration?

Service validation

Is the deployed service actually functioning?

Operational validation

Does the service recover and behave appropriately under expected operating conditions?

These checks become particularly important for DNS, monitoring, storage, and other supporting infrastructure services.

The automation environment continues to evolve.

Areas for further development include:

  • Improving consistency across host inventories
  • Expanding service validation
  • Strengthening change and rollback procedures
  • Documenting role dependencies
  • Improving maintenance reporting
  • Evaluating infrastructure-as-code workflows beyond traditional configuration management

Automation is most useful when it improves consistency, repeatability, maintainability, and recovery.

The goal is not simply to replace commands with YAML. It is to create a dependable method for managing infrastructure throughout its lifecycle.