Business Continuity Policy

Disaster recovery, failover procedures, RPO/RTO targets.

NIST controls:CP-1 through CP-10
Last updated:2026-03-15
Category:Governance
# Business Continuity Policy **Document ID**: KD-POL-005 **Owner**: Chief Technology Officer **Approved By**: CEO, Arivaran **Effective Date**: 2026-03-19 **Next Review Date**: 2027-03-19 (recomputed 2026-10-02 on the stated semi-annual cycle from the Effective Date; the 2026-09-19 occurrence was not recorded, and this date does not assert that a review has taken place) **Review Cycle**: Semi-annual **Classification**: Internal ## 1. Purpose This policy defines the business continuity and disaster recovery requirements for KuiperDesk. As a backup-as-a-service platform, KuiperDesk's availability is directly tied to customers' ability to recover their own data. A KuiperDesk outage during a customer disaster is a compounded failure. The continuity architecture must ensure that data durability survives infrastructure failures and that restore operations remain available even during partial platform degradation. ## 2. Scope This policy covers: - Availability and disaster recovery of all KuiperDesk backend services (aaHashSvc, aaTwinSvc, PostgreSQL (PgEdge on hub, CNPG on satellites), SeaweedFS, Keycloak) - Data durability of customer backup blocks and metadata - Network availability (WireGuard overlay, DNS, ingress) - Recovery of the CI/CD pipeline and management plane - Customer-facing restore operations during degraded states Out of scope: Customer endpoint availability (the aaagent can queue locally if the backend is unreachable). ## 3. RTO and RPO Targets | Service | RPO | RTO | Justification | |---------|-----|-----|---------------| | **SeaweedFS (backup blocks)** | 0 (replication) | 4 hours | Customer data store. 001 replication provides synchronous durability. | | **PostgreSQL (hash metadata)** | 0 (hub PgEdge multi-master) / 5 min (satellite CNPG WAL) | 0 (hub) / 2 hours (satellite) | Hub: PgEdge multi-master, both nodes writable. Satellite: CNPG streaming replication. | | **PostgreSQL (backup catalog)** | 0 (hub PgEdge) / 5 min (satellite CNPG WAL) | 0 (hub) / 2 hours (satellite) | Hub: instant failover. Satellite: CNPG HA with automatic failover. | | **Keycloak (identity)** | 1 hour (DB backup) | 4 hours | Backed by PgEdge (hub) or CNPG (satellite). Identity data changes infrequently. | | **aaHashSvc / aaTwinSvc** | N/A (stateless) | 30 minutes | Kubernetes restarts pods. State is in PostgreSQL (PgEdge on hub, CNPG on satellites). | | **CI/CD (Forgejo + Runner)** | 24 hours | 8 hours | Git repos are mirrored to GitHub. Runner is rebuildable. | | **Monitoring (VictoriaMetrics + Loki)** | 24 hours | 8 hours | Monitoring loss is operational, not customer-facing. | ## 4. Policy Statements ### 4.1 Data Durability Architecture 4.1.1. **Hub PostgreSQL (PgEdge Multi-Master)**: Hub hash metadata and backup catalog are stored in PostgreSQL 18 managed by PgEdge with Spock 5 async logical replication. Both hub nodes (EU1 + US1) are writable. Instant failover — no promotion needed. k8gb routes traffic to nearest hub. 4.1.2. **Satellite PostgreSQL (CNPG)**: Satellite hash metadata and backup catalog are stored in PostgreSQL/Citus managed by CloudNativePG (CNPG). CNPG provides streaming replication within the satellite for HA. Automatic failover promotes the standby within 30 seconds of primary failure detection. 4.1.3. **SeaweedFS (Backup Blocks)**: SeaweedFS is configured with 001 replication (one copy). Data durability for the single-node deployment relies on the underlying NVMe storage and the PostgreSQL continuous WAL archiving that captures volume metadata. In the event of total SeaweedFS data loss, customer backup blocks must be re-uploaded from agents (the agents retain local block copies until upload confirmation). 4.1.4. **Planned Enhancement**: SeaweedFS replication factor will be increased to 002 when a second storage node is provisioned. This is tracked as a risk item in the risk register. ### 4.2 Backup and Recovery Procedures 4.2.1. **Hub PostgreSQL Backup (PgEdge)**: PgEdge multi-master replication provides real-time data durability across both hub nodes. Additionally, pg_basebackup to SeaweedFS S3 runs daily. Point-in-time recovery (PITR) is available to any point within the last 7 days. 4.2.2. **Satellite PostgreSQL Backup (CNPG)**: CNPG continuous archiving streams WAL segments to SeaweedFS S3. A base backup is taken daily. Point-in-time recovery (PITR) is available to any point within the last 7 days. 4.2.3. **etcd Backup**: RKE2 etcd is backed up daily via the built-in snapshot mechanism. Snapshots are stored on the node filesystem and replicated to SeaweedFS. Recovery restores the Kubernetes cluster state. 4.2.4. **Forgejo Backup**: Git repositories are mirrored to GitHub (github.com/arivaran-ai/arivaran-ai) on every push. Forgejo database is backed up daily to SeaweedFS. 4.2.5. **Backup Verification**: Backup integrity is verified weekly by restoring to a test namespace and running a validation query (PostgreSQL: run a pg_dump of a sample schema; etcd: verify snapshot integrity with etcdctl). ### 4.3 Disaster Recovery Scenarios #### Scenario 1: Single Pod/Container Failure - **Detection**: Kubernetes liveness/readiness probes, VictoriaMetrics alerts - **Response**: Automatic. Kubernetes restarts the pod. No human intervention required. - **Recovery Time**: 30 seconds to 2 minutes #### Scenario 2: Single Node Failure (1 of 3 RKE2 nodes) - **Detection**: Node NotReady status, Gatus endpoint failure - **Response**: Kubernetes reschedules workloads to surviving nodes. Satellite CNPG loses one replica but maintains quorum. Hub PgEdge: other hub node continues serving. - **Recovery**: Replace the node via Ansible playbook (base-setup.yml + deploy-base.yml). Satellite CNPG re-replicates automatically. Hub PgEdge: Spock catches up automatically. - **Recovery Time**: 2-4 hours for full replica re-replication #### Scenario 3: Complete Loss of One Hub Node - **Detection**: Gatus and VPS-3 monitoring detect hub endpoint loss; k8gb detects node down (~30s) - **Response**: k8gb automatically routes traffic to the surviving hub node. PgEdge multi-master means the other hub is already writable — no promotion needed. - **Recovery Steps**: 1. k8gb DNS failover routes all traffic to surviving hub (automatic) 2. Verify PgEdge replication is healthy on surviving node 3. Re-provision failed node via Ansible (phase0 + phase1 + deploy-platform) 4. Re-join PgEdge replication mesh 5. Verify k8gb re-adds recovered node - **Recovery Time**: Automatic failover <1 minute; full re-provision 2-4 hours - **Data Loss**: Near-zero (PgEdge async replication lag, typically <1s) #### Scenario 4: OVH Provider-Wide Outage - **Detection**: Both hub nodes unreachable (different OVH DCs), OVH status page confirmation - **Response**: Satellites continue serving agents independently (data plane unaffected). Admin dashboard unavailable until at least one hub recovers or is re-provisioned at an alternative provider. - **Recovery Time**: Depends on OVH; re-provision at alternative provider via IaC: 4-8 hours ### 4.4 Hub HA Model 4.4.1. The product hub operates as an identical HA pair: EU1 (Milan) and US1 (Oregon). Both nodes run independent single-node RKE2 clusters with PgEdge multi-master PostgreSQL. Both are writable simultaneously — no primary/standby distinction for writes. 4.4.2. k8gb provides DNS-based geo-routing and automatic failover (~30s detection). Admin dashboard traffic routes to the nearest healthy hub. 4.4.3. PgEdge/Spock async logical replication keeps both hubs synchronized. On hub death, the other hub continues serving without promotion. 4.4.4. Both hubs are managed by unified IaC (Ansible + OpenTofu). Re-provisioning a failed hub is a standard playbook run, not a manual DR procedure. ### 4.5 DR Testing 4.5.1. DR failover is tested semi-annually. The test includes: k8gb failover verification (disable one hub), PgEdge replication catch-up verification, and end-to-end agent connectivity test. 4.5.2. Test results are documented with: recovery time achieved vs. RTO target, data loss measured vs. RPO target, and any issues encountered. 4.5.3. DR test failures generate P3 incidents and remediation action items. ### 4.6 Agent Resilience 4.6.1. The aaagent Windows client is designed to tolerate backend unavailability. When the backend is unreachable, the agent queues backup blocks locally and retries with exponential backoff. No customer backup data is lost due to a KuiperDesk backend outage. 4.6.2. Agents cache their last-known configuration and can continue performing local backups to the RocksDB block store during an extended outage. ## 5. Roles and Responsibilities | Role | Responsibility | |------|---------------| | CTO | Authorize DR failover. Define RTO/RPO targets. Approve DR test plans. | | Infrastructure Engineers | Execute DR procedures. Maintain backup CronJobs. Conduct DR tests. | | CEO | Approve customer communication during extended outages. | ## 6. Compliance Mapping | Policy Statement | SOC 2 Criteria | ISO 27001:2022 | |-----------------|----------------|----------------| | 4.1 Data Durability | A1.2 | A.8.13, A.8.14 | | 4.2 Backup Procedures | A1.2 | A.8.13 | | 4.3 DR Scenarios | A1.1, A1.3 | A.5.29, A.5.30 | | 4.5 DR Testing | A1.3 | A.5.30 | | 4.6 Agent Resilience | A1.1 | A.8.14 | ## 7. Exceptions Exceptions to RTO/RPO targets must be documented in customer SLAs. Any degradation of replication factors (e.g., running satellite CNPG without streaming replicas, or disabling PgEdge Spock replication on hub) requires CTO approval and a risk register entry. ## 8. Enforcement Backup CronJob failures generate automated alerts. Two consecutive backup failures for any system constitute a P3 incident. DR test schedule adherence is tracked and reported at quarterly risk reviews. ## 9. Revision History | Version | Date | Author | Changes | |---------|------|--------|---------| | 1.0 | 2026-03-19 | CTO | Initial release |

This document is part of the Arivaran Twin compliance program. For questions or the latest version, contact compliance@arivaran.ai.

Release-ready. Saved on this browser.