Satellite Database Architecture

PostgreSQL CNPG, Citus sharding, per-tenant isolation, backup strategy.

NIST controls:SC-28, CP-9, AC-3
Last updated:2026-03-22
Category:Technical
# Satellite Database Architecture — Single-Node Citus by Default ## Principle Every KuiperDesk deployment — hub or satellite — runs CNPG with Citus extension enabled. Schema is identical everywhere. The only difference is scale: single-node vs multi-node. ## Why Citus Everywhere (Even Single-Node) 1. **Schema portability** — `create_distributed_table('endpoints', 'tenant_id')` works on single-node Citus (data stays local, distribution is logical). No schema forks between hub and satellite. 2. **Seamless scale-up** — tenant outgrows single node → add CNPG instances → Citus redistributes shards. Zero migration, zero downtime. 3. **Query compatibility** — co-located joins, distributed aggregates, reference tables all work identically on 1 node or 10. 4. **Test-prod parity** — developers run single-node Citus locally, production runs multi-node. Same queries, same behavior. ## Satellite Tiers | Config | CNPG Instances | Citus Workers | Use Case | Endpoints | |--------|---------------|---------------|----------|-----------| | **Single-node** (default) | 1 | 0 (coordinator-only) | Small/medium Tier A, all Tier B/C | Up to ~5,000 | | **HA single-shard** | 3 (1 primary + 2 replica) | 0 | Tier A requiring HA | Up to ~5,000 | | **Multi-node** | 3+ | 2+ dedicated workers | Large Tier A enterprise | 5,000 — 100,000+ | | **Hub** | 3 | Scales with tenants | Central hub (all Tier B/C/D) | Unlimited | ## Database Layout Per Satellite ``` CNPG Cluster: kuiperdesk-db (Citus-enabled) │ ├─ Database: hashsvc │ Owner: hashsvc │ Tables: hash_metadata (distributed by tenant_id) │ Purpose: Content-addressable block dedup metadata │ ├─ Database: twinsvc │ Owner: twinsvc │ Tables: tenants, endpoints, snapshots, snapshot_disks, │ snapshot_volumes, cloud_actions, backup_policies, │ file_change_log, file_latest_state, synthetic_images, │ live_instances, twin_change_journals, pending_commands, │ agent_connections, regions, tenant_hierarchy │ All distributed by tenant_id (co-located) │ Purpose: Digital twin catalog, command channel, file index │ └─ Database: enrollsvc Owner: enrollsvc Tables: enrollment_keys, agent_versions, tenant_version_pins, tpm_challenges, computer_groups, computer_group_members, update_history, global_settings, proto_compatibility Purpose: Agent enrollment, version management ``` ## Citus Configuration ### Single-Node (Default) ```sql -- On satellite bootstrap, all distributed table calls succeed on single-node: CREATE EXTENSION IF NOT EXISTS citus; CREATE EXTENSION IF NOT EXISTS pg_trgm; -- These work on single-node — data stays local, distribution is logical SELECT create_distributed_table('endpoints', 'tenant_id'); SELECT create_distributed_table('snapshots', 'tenant_id', colocate_with => 'endpoints'); -- ... (all tables from 001_schema.sql) -- On a satellite, there's only ONE tenant_id, so all data is on one shard anyway. -- Citus overhead on single-node is negligible (< 1% CPU, < 10MB RAM). ``` ### Scale-Up to Multi-Node When a satellite needs horizontal scaling: ```bash # 1. Increase CNPG instances (Ansible/Helm) helm upgrade kuiperdesk-satellite ... --set cnpg.instances=5 # 2. Add Citus workers (via citus_add_node on coordinator) SELECT citus_add_node('kuiperdesk-db-2.kuiperdesk-db.kd-db.svc', 5432); SELECT citus_add_node('kuiperdesk-db-3.kuiperdesk-db.kd-db.svc', 5432); # 3. Rebalance shards across new workers SELECT rebalance_table_shards(); # Zero downtime. No schema changes. Existing queries work unchanged. ``` ## Satellite-to-Hub Data Flow ``` Satellite (tenant's infra) Hub (Arivaran infra) ┌──────────────────────────┐ ┌──────────────────────────┐ │ CNPG+Citus (kuiperdesk-db)│ │ CNPG+Citus (kuiperdesk-db)│ │ hashsvc DB │ │ hashsvc DB │ │ twinsvc DB │ │ twinsvc DB │ │ enrollsvc DB │ │ enrollsvc DB │ │ │ │ │ │ SeaweedFS (blocks + DBs) │ │ │ │ │ │ │ │ Agent ←→ Local Services │ │ Dashboard (aaDash) │ └──────────┬───────────────┘ └──────────┬───────────────┘ │ │ │ satgw tunnel (TLS 1.3 mTLS) │ │ │ └── Metadata sync ──────────────────────┘ - Satellite status heartbeat - Aggregated backup metrics - Billing metering data - Alert escalation NOT synced: - Backup blocks (stay on satellite S3) - File indexes (stay on satellite PG) - Agent credentials (stay on satellite) ``` ## Data Sovereignty | Data Type | Location | Leaves Satellite? | |-----------|----------|-------------------| | Backup blocks (S3) | Satellite SeaweedFS / AWS S3 | **Never** | | Backup SQLite DBs | Satellite S3 | **Never** | | Hash metadata | Satellite PostgreSQL | **Never** | | File indexes | Satellite PostgreSQL | **Never** | | Agent credentials | Satellite (TPM + config.db) | **Never** | | Endpoint list | Satellite PostgreSQL | Summary count only (metering) | | Backup status | Satellite PostgreSQL | Last backup time only (dashboard) | | Alert events | Satellite → Hub | Yes (for centralized alerting) | | Billing metrics | Satellite → Hub | Yes (aggregated storage/endpoint counts) | This ensures compliance with data residency requirements (GDPR, sovereign cloud, FedRAMP boundary). ## Backup & Recovery ### Satellite Database Backup Each satellite's CNPG cluster is configured with: - **WAL archiving** to local SeaweedFS S3 (via barman-cloud) - **Scheduled base backups** daily at 03:00 local time - **PITR** (Point-in-Time Recovery) to any second within retention window - **Retention**: 30 days (configurable per tenant) ### Disaster Recovery If a satellite is destroyed: 1. New satellite bootstrapped (same bootstrap token flow) 2. CNPG cluster bootstrapped from S3 backup (barman-cloud recovery) 3. Agents re-enroll (enrollment keys still valid on hub) 4. Data restored from S3 (blocks are immutable, SQLite DBs intact) If satellite S3 is also destroyed: - Hub has no copy of the data (by design — data sovereignty) - Tenant must restore from their own offsite backup - Recommendation: configure S3 cross-region replication for Tier A ## Compliance Mapping | Framework | Control | How This Architecture Satisfies | |-----------|---------|-------------------------------| | **SOC 2** CC6.5 | Data classification | Backup data classified and contained within satellite boundary | | **FedRAMP** SC-28 | Protection of information at rest | Citus + CNPG with encrypted storage (LUKS/EBS encryption) | | **FedRAMP** CP-9 | Information system backup | Automated WAL archiving + daily base backups to S3 | | **GDPR** Art. 44-49 | International transfers | Data never leaves satellite geography (customer controls location) | | **HIPAA** §164.310(d) | Device and media controls | Database on customer-controlled hardware with encryption at rest | ## Revision History | Date | Change | Author | |------|--------|--------| | 2026-03-22 | Initial architecture — single-node Citus default, scale-up path, data sovereignty model | Lena Nguyen (Backend Architect), Daniel Okafor (Cloud Infrastructure) |

This document is part of the Arivaran Twin compliance program. For questions or the latest version, contact compliance@arivaran.ai.

Release-ready. Saved on this browser.