Skip to content

Aesir Production Cluster

Aesir is the primary production virtualization environment within OnyxNet.

It consists of three similarly configured HP EliteDesk 800 G5 Mini systems running Proxmox VE.

The cluster provides an environment for hosting self-managed applications and infrastructure services while exploring clustered virtualization, resource management, and recovery planning.

Status: Operational

Role: Primary production virtualization cluster

Physical nodes: Three

Development direction: Existing deployment under review

The cluster is operational, but parts of its workload and resource allocation strategy are being reevaluated.

This documentation does not imply that every proposed optimization has already been implemented.

The three physical nodes have the same base hardware configuration.

Component Per-node configuration
System HP EliteDesk 800 G5 Mini
Processor Intel Core i5-9500T
Memory 16 GB DDR4
Storage 2 × Patriot P300 128 GB NVMe
Network 1 GbE Intel NIC
Graphics Intel integrated graphics

Across the three nodes, the cluster contains:

  • Three physical processors
  • 48 GB of installed system memory
  • Six 128 GB NVMe drives

These totals describe installed hardware, not the resources available to individual virtual machines.

Proxmox, storage, and other host-level functions also consume system resources.

Aesir uses Proxmox clustering to manage the three physical nodes as one administrative environment.

The nodes are designed to support distributed workload placement and operational flexibility.

Cluster membership, storage replication, and automatic high availability are separate concepts.

The presence of multiple nodes alone does not guarantee automatic workload recovery after a hardware failure.

The existing design includes:

  • Proxmox local storage
  • Local LVM-thin storage
  • Node-local ZFS storage

The ZFS pool named vm-ct_storage exists on each Aesir node.

These pools are separate local storage systems, not one shared storage pool.

This distinction is important when evaluating migration, replication, and availability.

A guest using node-local storage cannot assume that its disk is immediately accessible from another node.

Depending on the chosen procedure and configuration, moving or recovering a guest may involve:

  • Transferring virtual disks
  • Restoring a backup
  • Using appropriately configured storage replication
  • Rebuilding a service from documented configuration

Any availability or recovery claim must reflect the mechanism actually deployed and tested.

Aesir is intended for services that support the normal operation of the homelab.

Potential workload categories include:

Category Examples
Infrastructure Reverse proxy and supporting services
Applications Self-hosted applications
Observability Monitoring and metrics collection
Operations Management and administration
On-demand workloads Services that need not run continuously

Actual service placement and runtime status may change during the cluster optimization work.

The current workload inventory should be revalidated before publishing detailed per-node assignments.

Each node uses an existing 1 GbE network interface.

The cluster operates within OnyxNet’s broader segmented network architecture.

The virtualization design must account for network capacity, host availability, and service connectivity.

A future faster-networking upgrade is a long-term objective, not a current requirement.

The homelab includes Proxmox Backup Server.

Production workload protection should be evaluated using several independent questions:

  1. Is the guest included in scheduled backups?
  2. Are backup jobs completing successfully?
  3. Is retention configured appropriately?
  4. Can the backup be restored successfully?
  5. Are service dependencies documented?
  6. Does recovery meet the workload’s requirements?

Successful backup creation alone does not prove that an application can be recovered within an acceptable time.

The current architecture must work within existing hardware limitations.

Primary constraints include:

  • 16 GB RAM per physical node
  • Limited local NVMe storage capacity
  • 1 GbE host networking
  • Node-local rather than shared ZFS storage
  • Avoiding unnecessary service disruption
  • No immediate budget for hardware expansion

These constraints directly influence workload selection, service consolidation, and resource allocation decisions.

Future improvements may include:

  • Updating and validating the workload inventory
  • Reviewing resource allocations
  • Reconsidering guest placement
  • Removing unnecessary or duplicate services
  • Improving backup and recovery documentation
  • Evaluating appropriate replication mechanisms
  • Defining measurable operational requirements

These are review areas rather than claims of completed work.

A production virtualization environment should be designed around operational needs, not merely the number of guests it can host.

Resource constraints, storage locality, recoverability, and maintainability are central architectural considerations.