# Satellite Database Architecture — Single-Node Citus by Default
## Principle
Every KuiperDesk deployment — hub or satellite — runs CNPG with Citus extension enabled. Schema is identical everywhere. The only difference is scale: single-node vs multi-node.
## Why Citus Everywhere (Even Single-Node)
1. **Schema portability** — `create_distributed_table('endpoints', 'tenant_id')` works on single-node Citus (data stays local, distribution is logical). No schema forks between hub and satellite.
2. **Seamless scale-up** — tenant outgrows single node → add CNPG instances → Citus redistributes shards. Zero migration, zero downtime.
3. **Query compatibility** — co-located joins, distributed aggregates, reference tables all work identically on 1 node or 10.
4. **Test-prod parity** — developers run single-node Citus locally, production runs multi-node. Same queries, same behavior.
## Satellite Tiers
| Config | CNPG Instances | Citus Workers | Use Case | Endpoints |
|--------|---------------|---------------|----------|-----------|
| **Single-node** (default) | 1 | 0 (coordinator-only) | Small/medium Tier A, all Tier B/C | Up to ~5,000 |
| **HA single-shard** | 3 (1 primary + 2 replica) | 0 | Tier A requiring HA | Up to ~5,000 |
| **Multi-node** | 3+ | 2+ dedicated workers | Large Tier A enterprise | 5,000 — 100,000+ |
| **Hub** | 3 | Scales with tenants | Central hub (all Tier B/C/D) | Unlimited |
## Database Layout Per Satellite
```
CNPG Cluster: kuiperdesk-db (Citus-enabled)
│
├─ Database: hashsvc
│ Owner: hashsvc
│ Tables: hash_metadata (distributed by tenant_id)
│ Purpose: Content-addressable block dedup metadata
│
├─ Database: twinsvc
│ Owner: twinsvc
│ Tables: tenants, endpoints, snapshots, snapshot_disks,
│ snapshot_volumes, cloud_actions, backup_policies,
│ file_change_log, file_latest_state, synthetic_images,
│ live_instances, twin_change_journals, pending_commands,
│ agent_connections, regions, tenant_hierarchy
│ All distributed by tenant_id (co-located)
│ Purpose: Digital twin catalog, command channel, file index
│
└─ Database: enrollsvc
Owner: enrollsvc
Tables: enrollment_keys, agent_versions, tenant_version_pins,
tpm_challenges, computer_groups, computer_group_members,
update_history, global_settings, proto_compatibility
Purpose: Agent enrollment, version management
```
## Citus Configuration
### Single-Node (Default)
```sql
-- On satellite bootstrap, all distributed table calls succeed on single-node:
CREATE EXTENSION IF NOT EXISTS citus;
CREATE EXTENSION IF NOT EXISTS pg_trgm;
-- These work on single-node — data stays local, distribution is logical
SELECT create_distributed_table('endpoints', 'tenant_id');
SELECT create_distributed_table('snapshots', 'tenant_id', colocate_with => 'endpoints');
-- ... (all tables from 001_schema.sql)
-- On a satellite, there's only ONE tenant_id, so all data is on one shard anyway.
-- Citus overhead on single-node is negligible (< 1% CPU, < 10MB RAM).
```
### Scale-Up to Multi-Node
When a satellite needs horizontal scaling:
```bash
# 1. Increase CNPG instances (Ansible/Helm)
helm upgrade kuiperdesk-satellite ... --set cnpg.instances=5
# 2. Add Citus workers (via citus_add_node on coordinator)
SELECT citus_add_node('kuiperdesk-db-2.kuiperdesk-db.kd-db.svc', 5432);
SELECT citus_add_node('kuiperdesk-db-3.kuiperdesk-db.kd-db.svc', 5432);
# 3. Rebalance shards across new workers
SELECT rebalance_table_shards();
# Zero downtime. No schema changes. Existing queries work unchanged.
```
## Satellite-to-Hub Data Flow
```
Satellite (tenant's infra) Hub (Arivaran infra)
┌──────────────────────────┐ ┌──────────────────────────┐
│ CNPG+Citus (kuiperdesk-db)│ │ CNPG+Citus (kuiperdesk-db)│
│ hashsvc DB │ │ hashsvc DB │
│ twinsvc DB │ │ twinsvc DB │
│ enrollsvc DB │ │ enrollsvc DB │
│ │ │ │
│ SeaweedFS (blocks + DBs) │ │ │
│ │ │ │
│ Agent ←→ Local Services │ │ Dashboard (aaDash) │
└──────────┬───────────────┘ └──────────┬───────────────┘
│ │
│ satgw tunnel (TLS 1.3 mTLS) │
│ │
└── Metadata sync ──────────────────────┘
- Satellite status heartbeat
- Aggregated backup metrics
- Billing metering data
- Alert escalation
NOT synced:
- Backup blocks (stay on satellite S3)
- File indexes (stay on satellite PG)
- Agent credentials (stay on satellite)
```
## Data Sovereignty
| Data Type | Location | Leaves Satellite? |
|-----------|----------|-------------------|
| Backup blocks (S3) | Satellite SeaweedFS / AWS S3 | **Never** |
| Backup SQLite DBs | Satellite S3 | **Never** |
| Hash metadata | Satellite PostgreSQL | **Never** |
| File indexes | Satellite PostgreSQL | **Never** |
| Agent credentials | Satellite (TPM + config.db) | **Never** |
| Endpoint list | Satellite PostgreSQL | Summary count only (metering) |
| Backup status | Satellite PostgreSQL | Last backup time only (dashboard) |
| Alert events | Satellite → Hub | Yes (for centralized alerting) |
| Billing metrics | Satellite → Hub | Yes (aggregated storage/endpoint counts) |
This ensures compliance with data residency requirements (GDPR, sovereign cloud, FedRAMP boundary).
## Backup & Recovery
### Satellite Database Backup
Each satellite's CNPG cluster is configured with:
- **WAL archiving** to local SeaweedFS S3 (via barman-cloud)
- **Scheduled base backups** daily at 03:00 local time
- **PITR** (Point-in-Time Recovery) to any second within retention window
- **Retention**: 30 days (configurable per tenant)
### Disaster Recovery
If a satellite is destroyed:
1. New satellite bootstrapped (same bootstrap token flow)
2. CNPG cluster bootstrapped from S3 backup (barman-cloud recovery)
3. Agents re-enroll (enrollment keys still valid on hub)
4. Data restored from S3 (blocks are immutable, SQLite DBs intact)
If satellite S3 is also destroyed:
- Hub has no copy of the data (by design — data sovereignty)
- Tenant must restore from their own offsite backup
- Recommendation: configure S3 cross-region replication for Tier A
## Compliance Mapping
| Framework | Control | How This Architecture Satisfies |
|-----------|---------|-------------------------------|
| **SOC 2** CC6.5 | Data classification | Backup data classified and contained within satellite boundary |
| **FedRAMP** SC-28 | Protection of information at rest | Citus + CNPG with encrypted storage (LUKS/EBS encryption) |
| **FedRAMP** CP-9 | Information system backup | Automated WAL archiving + daily base backups to S3 |
| **GDPR** Art. 44-49 | International transfers | Data never leaves satellite geography (customer controls location) |
| **HIPAA** §164.310(d) | Device and media controls | Database on customer-controlled hardware with encryption at rest |
## Revision History
| Date | Change | Author |
|------|--------|--------|
| 2026-03-22 | Initial architecture — single-node Citus default, scale-up path, data sovereignty model | Lena Nguyen (Backend Architect), Daniel Okafor (Cloud Infrastructure) |