Backup & Recovery Strategy
Backups are an essential part of infrastructure operations, but reliable recovery requires more than saving copies of data.
OnyxNet’s recovery strategy focuses on preserving important information, understanding failure scenarios, and establishing repeatable restoration procedures.
This document outlines the recovery principles used to evaluate the existing environment.
It does not claim that every described control has already been implemented.
Implementation Status
Section titled “Implementation Status”Current backup platforms:
- TrueNAS with ZFS
- Proxmox Backup Server
Offsite disaster recovery: Deferred
Recovery testing: Formal documentation pending
The immediate objective is to document existing coverage and verify which recovery processes are already functional.
Protection Layers
Section titled “Protection Layers”Several technologies contribute to resilience, but they address different problems.
| Protection | Primary Purpose |
|---|---|
| Disk redundancy | Tolerate certain hardware failures |
| ZFS snapshots | Recover earlier dataset states |
| PBS backups | Recover Proxmox guest data |
| Independent backup copies | Reduce dependence on the source system |
| Offsite backups | Protect against site-level disasters |
| Restore testing | Demonstrate recovery capability |
These layers are complementary rather than interchangeable.
Recovery Objectives
Section titled “Recovery Objectives”Two concepts help define operational recovery expectations.
Recovery Point Objective (RPO)
Section titled “Recovery Point Objective (RPO)”RPO describes the maximum acceptable data loss measured in time.
For example, a workload with an RPO of 24 hours may tolerate losing changes made after its most recent daily recovery point.
Recovery Time Objective (RTO)
Section titled “Recovery Time Objective (RTO)”RTO describes the target time to restore an affected service.
A workload may have a valid backup but still take hours to restore.
RPO and RTO must be selected according to the operational importance of the workload.
OnyxNet has not yet published measured RPO or RTO values for individual services.
Failure Scenarios
Section titled “Failure Scenarios”A useful recovery strategy should consider multiple classes of failure.
Accidental Deletion
Section titled “Accidental Deletion”Possible recovery sources include earlier ZFS snapshots or independent backups.
The available options depend on dataset coverage and retention settings.
Application Failure
Section titled “Application Failure”Restoring an application may require:
- Application files
- Configuration
- Database content
- Supporting services
- Network connectivity
- Credentials or encryption keys
Restoring the virtual machine alone may not restore the full service.
Virtualization Host Failure
Section titled “Virtualization Host Failure”A failed Proxmox node may require recovery onto another compatible host.
Storage accessibility, available resources, and backup integrity directly affect recovery options.
Storage Failure
Section titled “Storage Failure”Loss of a storage device or pool may require restoring data from an independent backup.
Local snapshots do not protect against every storage failure.
Site-Level Disaster
Section titled “Site-Level Disaster”A fire, theft, environmental incident, or major electrical event could affect multiple local systems simultaneously.
An independent offsite copy is a future resilience goal.
That capability is not currently claimed as implemented.
Backup Verification vs. Restore Testing
Section titled “Backup Verification vs. Restore Testing”Backup integrity verification and application recovery tests serve different purposes.
Backup Verification
Section titled “Backup Verification”Verification checks whether stored backup data meets the backup system’s integrity requirements.
For Proxmox Backup Server, verification jobs are one mechanism for assessing stored backup integrity.
Restore Testing
Section titled “Restore Testing”Restore testing exercises the process of recovering data or a workload.
A successful test should confirm:
- The intended recovery point exists.
- The data can be retrieved.
- The recovery destination is suitable.
- Required files or services are usable.
- The procedure can be repeated.
- Results and limitations are documented.
A verification job does not replace a real restore test.
Recommended Recovery Test Categories
Section titled “Recommended Recovery Test Categories”File Recovery
Section titled “File Recovery”Restore a selected file or small dataset to a separate location.
Confirm file integrity and usability.
Virtual Machine Recovery
Section titled “Virtual Machine Recovery”Restore a noncritical guest into a controlled recovery environment.
Avoid network or identity conflicts with the production guest.
Linux Container Recovery
Section titled “Linux Container Recovery”Restore a container from a selected backup and validate the application and supporting dependencies.
Service Recovery
Section titled “Service Recovery”Verify that a recovered service is functionally usable rather than simply powered on.
For example, a restored application may still depend on databases, DNS, network routes, or external storage.
Recovery Test Record
Section titled “Recovery Test Record”A standardized record helps make testing results useful over time.
| Field | Description |
|---|---|
| Test date | When the test occurred |
| Service | What was recovered |
| Failure scenario | What the test simulated |
| Recovery source | Backup or snapshot used |
| Recovery destination | Where data was restored |
| Result | Passed, failed, or partial |
| Recovery duration | Measured elapsed time |
| Findings | Problems and lessons learned |
Do not publish credentials, private host addresses, or sensitive recovery instructions in public test records.
Retention and Capacity
Section titled “Retention and Capacity”Backup retention should balance:
- Required recovery history
- Available storage
- Backup frequency
- Application change rate
- Regulatory or personal requirements
- Recovery objectives
More recovery points are not automatically better if they exhaust available storage or prevent newer backups from completing.
For PBS, pruning and garbage collection are different operations and should be understood before changing retention or storage-reclamation settings.
Offsite Backup Roadmap
Section titled “Offsite Backup Roadmap”The long-term Ginnungagap project aims to maintain an independent backup system at another location.
The currently selected approach is:
- TrueNAS SCALE for the remote appliance
- Tailscale for encrypted connectivity
- ZFS replication for selected NAS data
- A suitable PBS-compatible mechanism for Proxmox VM/LXC backups
The project is deferred due to budget constraints and is not currently deployed.
VLAN 700 remains documented as a superseded historical design proposal.
Current Improvement Opportunities
Section titled “Current Improvement Opportunities”Before investing in additional hardware, useful operational improvements include:
- Inventorying protected workloads.
- Confirming existing backup schedules.
- Reviewing snapshot and retention settings.
- Checking the latest backup results.
- Testing representative restorations.
- Identifying unprotected dependencies.
- Documenting recovery procedures.
- Monitoring backup failures and capacity.
The purpose is to improve confidence in the systems already available.
Engineering Takeaway
Section titled “Engineering Takeaway”A completed backup job is evidence that data was written.
A successful restore test is stronger evidence that the data can be recovered.
Dependable infrastructure requires both protection and tested recovery processes.