This document establishes standard operating procedures (SOPs), maintenance schedules, emergency failover protocols, hardware-triggered telemetry responses, and diagnostic runbooks for the stefanutc1/infrastructure platform.
It equips Systems Administrators, Site Reliability Engineers (SREs), and Platform Engineers with deterministic workflows for both day-to-day operations and disaster scenarios.
| Frequency | Task / Procedure | Automation Script / Tool | Expected Duration |
|---|---|---|---|
| Continuous (Real-Time) | Endpoint Availability & Latency Monitoring | Uptime Kuma (CT 103) & Prometheus (CT 104) | Automated (<15s scrape) |
| Continuous (Real-Time) | Edge Telemetry, Rack Thermals & Grid Watch | ESP32 Nodes 01–04 (esp32/) & Home Assistant |
Sub-second edge interrupts |
| Daily (02:00 UTC) | vzdump Container & VM Snapshots to OMV NAS | Proxmox vzdump Engine & PBS Client | ~25 minutes |
| Daily (00:00 Europe/Bucharest) | Threat Intelligence & Currency Sync | GitHub Actions CD & Python Sync Scripts | ~3 minutes |
| Weekly (Sundays 03:00 UTC) | Operating System Package Updates (Debian/PVE) | ansible/playbooks/maintenance.yml |
~15 minutes |
| Weekly (Sundays 04:00 UTC) | ZFS Storage Pool Scrub & Health Audit | OpenMediaVault ZFS Engine (zpool scrub) |
~45 minutes |
| Monthly | Backup Restoration Verification Drill | docs/runbooks/vzdump_restore_drill.md |
~20 minutes |
| Quarterly | WireGuard VPN & TLS Root Key Rotation | scripts/wireguard_key_rotation.sh |
~30 minutes |
Restores full datacenter operations from a cold, unpowered state in strict dependency-ordered stages: - **Runbook**: [`docs/runbooks/cold_boot_sequence.md`](docs/runbooks/cold_boot_sequence.md) - **Automated Script**: `scripts/cold-boot-sequence.sh` (or `scripts/cold-boot-sequence.ps1`) - **Key Verification Gate**: Confirm OPNsense (VM 200) is running and resolving DNS before booting downstream LXC containers. Gracefully terminates active databases, containers, and hypervisors during power loss or thermal alarms: - **Runbook**: [`docs/runbooks/emergency_shutdown.md`](docs/runbooks/emergency_shutdown.md) - **Automated Script**: `scripts/emergency-shutdown.sh` (or `scripts/emergency-shutdown.ps1`) - **Safety Guarantee**: Flushes ZFS transaction groups and syncs NVMe journal pages to prevent filesystem corruption. - **Node**: `ESP32-EDGE-04` (`esp32/power_monitor/`) - **Mechanism**: The 230V AC optocoupler detects grid drop instantly. If power remains lost and the 12V SLA battery bank discharges below **11.4V (critical cut-off threshold)**: 1. ESP32 sounds the onboard emergency buzzer. 2. Issues an authenticated emergency webhook to Proxmox VE host API (`POST /api2/json/nodes/pve/status -d command=shutdown`). 3. Proxmox automatically initiates `scripts/emergency-shutdown.sh`, safely powering down VMs, flushing databases, and unmounting NFS before battery cutoff. - **Node**: `ESP32-EDGE-03` (`esp32/datacenter_environment/`) - **Mechanism**: Continuously samples BME280 ambient temperature and dual DS18B20 1-Wire intake/exhaust probes: - If delta-T ($\Delta T = T_{\text{exhaust}} - T_{\text{intake}}$) exceeds $8.0^\circ\text{C}$ or exhaust exceeds $38.0^\circ\text{C}$, the 25kHz PWM driver spins Noctua cooling fans to 100% duty cycle. - If ambient temperature exceeds $45.0^\circ\text{C}$ despite fan cooling, an alert is dispatched to Prometheus Alertmanager and Uptime Kuma. Validates that backup archives are valid and bootable without causing production network conflicts: - **Runbook**: [`docs/runbooks/vzdump_restore_drill.md`](docs/runbooks/vzdump_restore_drill.md) - **Script**: `scripts/disaster-recovery/dr_vzdump_restore.sh` - **Method**: Restores backup to a temporary ID (900-series) on an isolated bridge (`vmbr3`), verifies process tables, and purges test artifacts.
Execute the master automated health validation script: ```bash python3 scripts/audit_infrastructure.py ``` This diagnostic engine inspects 12 operational domains (IaC, Ansible, Security, Secrets, Networking, Kubernetes, Observability, Backup, Disaster Recovery, Documentation, AI Governance, Supply Chain) and outputs an executive validation report. Verify compile integrity and configuration consistency across all 4 edge microcontroller sketches: ```bash python3 scripts/verify_esp32_firmware.py ``` Verify that all cloud Terraform configurations conform strictly to $0.00 / free-tier rules: ```bash python3 scripts/verify_zero_cloud_cost.py ``` To check live hypervisor telemetry and container statuses from the Proxmox console: ```bash bash scripts/healthcheck-fleet.sh ``` ```bash # Verify inter-firewall transit bus responsiveness: ping -c 2 10.10.20.1
dig @192.168.1.134 immich.lan +short
wg show wg-cloud0
---
<div align="center">
## 4. Maintenance & Rolling Updates
</div>
<div align="center">
### 4.1 Applying Operating System Updates
</div>
To perform safe, rolling updates across the container fleet using Ansible:
```bash
cd ansible
ansible-playbook -i inventories/homelab/hosts.yml playbooks/maintenance.yml
If root filesystem usage on Node 1 exceeds 80%: ```bash # Clean up downloaded package caches: apt-get clean # Clean up temporary vzdump cache: rm -rf /var/tmp/vzdump* # Trim LVM-thin pool: fstrim -av ``` Initiate monthly data scrub to verify block checksums: ```bash ssh root@192.168.1.135 "zpool scrub omv_tank" # Check scrub progress: ssh root@192.168.1.135 "zpool status omv_tank" ```
1. Identify offending process via `htop` or `top`. 2. If caused by local Ollama AI GPU inference, verify fan curves and thermal throttle. 3. If temperature continues rising past 85°C, initiate `docs/runbooks/emergency_shutdown.md`. 1. Check if backend container is running: `pct status `. 2. Inspect container systemd logs: `pct exec -- journalctl -xeu --no-pager -n 50`. 3. Check Caddy reverse proxy upstream logs: `docker logs caddy --tail 50`.
Engineered with precision by Moană Ștefănuț-Cornel (@stefanutc1).
Universitatea din Craiova · Facultatea de Economie și Administrarea Afacerilor (FEAA) · Informatică Economică (2024–2027).