Skip to content

Backup & Recovery Strategy

Backups are an essential part of infrastructure operations, but reliable recovery requires more than saving copies of data.

OnyxNet’s recovery strategy focuses on preserving important information, understanding failure scenarios, and establishing repeatable restoration procedures.

This document outlines the recovery principles used to evaluate the existing environment.

It does not claim that every described control has already been implemented.

Current backup platforms:

  • TrueNAS with ZFS
  • Proxmox Backup Server

Offsite disaster recovery: Deferred

Recovery testing: Formal documentation pending

The immediate objective is to document existing coverage and verify which recovery processes are already functional.

Several technologies contribute to resilience, but they address different problems.

Protection Primary Purpose
Disk redundancy Tolerate certain hardware failures
ZFS snapshots Recover earlier dataset states
PBS backups Recover Proxmox guest data
Independent backup copies Reduce dependence on the source system
Offsite backups Protect against site-level disasters
Restore testing Demonstrate recovery capability

These layers are complementary rather than interchangeable.

Two concepts help define operational recovery expectations.

RPO describes the maximum acceptable data loss measured in time.

For example, a workload with an RPO of 24 hours may tolerate losing changes made after its most recent daily recovery point.

RTO describes the target time to restore an affected service.

A workload may have a valid backup but still take hours to restore.

RPO and RTO must be selected according to the operational importance of the workload.

OnyxNet has not yet published measured RPO or RTO values for individual services.

A useful recovery strategy should consider multiple classes of failure.

Possible recovery sources include earlier ZFS snapshots or independent backups.

The available options depend on dataset coverage and retention settings.

Restoring an application may require:

  • Application files
  • Configuration
  • Database content
  • Supporting services
  • Network connectivity
  • Credentials or encryption keys

Restoring the virtual machine alone may not restore the full service.

A failed Proxmox node may require recovery onto another compatible host.

Storage accessibility, available resources, and backup integrity directly affect recovery options.

Loss of a storage device or pool may require restoring data from an independent backup.

Local snapshots do not protect against every storage failure.

A fire, theft, environmental incident, or major electrical event could affect multiple local systems simultaneously.

An independent offsite copy is a future resilience goal.

That capability is not currently claimed as implemented.

Backup integrity verification and application recovery tests serve different purposes.

Verification checks whether stored backup data meets the backup system’s integrity requirements.

For Proxmox Backup Server, verification jobs are one mechanism for assessing stored backup integrity.

Restore testing exercises the process of recovering data or a workload.

A successful test should confirm:

  1. The intended recovery point exists.
  2. The data can be retrieved.
  3. The recovery destination is suitable.
  4. Required files or services are usable.
  5. The procedure can be repeated.
  6. Results and limitations are documented.

A verification job does not replace a real restore test.

Restore a selected file or small dataset to a separate location.

Confirm file integrity and usability.

Restore a noncritical guest into a controlled recovery environment.

Avoid network or identity conflicts with the production guest.

Restore a container from a selected backup and validate the application and supporting dependencies.

Verify that a recovered service is functionally usable rather than simply powered on.

For example, a restored application may still depend on databases, DNS, network routes, or external storage.

A standardized record helps make testing results useful over time.

Field Description
Test date When the test occurred
Service What was recovered
Failure scenario What the test simulated
Recovery source Backup or snapshot used
Recovery destination Where data was restored
Result Passed, failed, or partial
Recovery duration Measured elapsed time
Findings Problems and lessons learned

Do not publish credentials, private host addresses, or sensitive recovery instructions in public test records.

Backup retention should balance:

  • Required recovery history
  • Available storage
  • Backup frequency
  • Application change rate
  • Regulatory or personal requirements
  • Recovery objectives

More recovery points are not automatically better if they exhaust available storage or prevent newer backups from completing.

For PBS, pruning and garbage collection are different operations and should be understood before changing retention or storage-reclamation settings.

The long-term Ginnungagap project aims to maintain an independent backup system at another location.

The currently selected approach is:

  • TrueNAS SCALE for the remote appliance
  • Tailscale for encrypted connectivity
  • ZFS replication for selected NAS data
  • A suitable PBS-compatible mechanism for Proxmox VM/LXC backups

The project is deferred due to budget constraints and is not currently deployed.

VLAN 700 remains documented as a superseded historical design proposal.

Before investing in additional hardware, useful operational improvements include:

  1. Inventorying protected workloads.
  2. Confirming existing backup schedules.
  3. Reviewing snapshot and retention settings.
  4. Checking the latest backup results.
  5. Testing representative restorations.
  6. Identifying unprotected dependencies.
  7. Documenting recovery procedures.
  8. Monitoring backup failures and capacity.

The purpose is to improve confidence in the systems already available.

A completed backup job is evidence that data was written.

A successful restore test is stronger evidence that the data can be recovered.

Dependable infrastructure requires both protection and tested recovery processes.