# Satellite Storage Architecture — Sensible Defaults by Scale
## Storage Backend Selection
| Citus Mode | Default Storage | Why | Override |
|------------|----------------|-----|----------|
| **single** (1 node, ≤5K endpoints) | SeaweedFS | Simple, single-binary, low overhead, S3-compatible | Can choose CubeFS |
| **ha** (3 nodes, ≤5K endpoints) | SeaweedFS (replicated) | SeaweedFS handles replication natively across volumes | Can choose CubeFS |
| **multi-node** (3+ nodes, 5K+ endpoints) | **CubeFS** | Distributed POSIX + S3, scales horizontally with nodes, erasure coding | Can choose SeaweedFS |
## Why CubeFS for Multi-Node
When a satellite scales to multi-node Citus:
- **Citus workers need shared storage** for WAL archiving, backup, and data redistribution
- **CubeFS** provides a distributed filesystem that grows with the cluster (add nodes → add capacity)
- **Erasure coding** reduces storage overhead vs 3x replication (1.5x overhead for same durability)
- **POSIX + S3 dual interface** — PostgreSQL WAL uses POSIX mount, agent uploads use S3 API
- **CSI driver** — native K8s StorageClass, PVCs work transparently
## Why SeaweedFS for Single-Node
- **Lower resource footprint** — master + volume + filer + s3 in ~500MB RAM total
- **Simple operations** — single-node, no distributed consensus
- **S3-compatible** — agent uploads work identically
- **Already proven** — QA Test #1 validated the full pipeline
## CubeFS Configuration (Multi-Node Default)
```yaml
# values-multi-node.yaml
storage:
backend: cubefs
cubefs:
enabled: true
master:
replicas: 3 # must match cluster node count
storage: 10Gi
metanode:
replicas: 3
storage: 20Gi
datanode:
replicas: 3
storage: 100Gi # bulk block storage
erasureCoding: true # 1.5x overhead vs 3x replication
objectnode:
replicas: 2 # S3 API gateway
port: 8333 # same port as SeaweedFS for transparent swap
csi:
enabled: true
storageClass: cubefs-sc
seaweedfs:
enabled: false
```
## Migration Path: SeaweedFS → CubeFS
When scaling from single-node to multi-node:
1. Deploy CubeFS alongside SeaweedFS (both running temporarily)
2. Migrate data: `aws s3 sync --endpoint-url <seaweedfs> s3://blocks/ --endpoint-url <cubefs> s3://blocks/`
3. Update S3 proxy to point to CubeFS objectnode
4. Verify agent uploads work against CubeFS
5. Decommission SeaweedFS
6. Total downtime: 0 (agent retries handle the cutover)
## Compliance Note
Both SeaweedFS and CubeFS support encryption at rest (LUKS on the underlying volumes). For FedRAMP:
- **SeaweedFS**: LUKS + volume-level encryption
- **CubeFS**: LUKS + erasure coding (data fragments spread across nodes, no single node has complete data)
CubeFS erasure coding provides additional security-through-distribution — a stolen disk contains only erasure-coded fragments, not complete files.
## Revision History
| Date | Change | Author |
|------|--------|--------|
| 2026-03-22 | Initial — CubeFS default for multi-node, SeaweedFS for single | Daniel Okafor (Cloud Infrastructure), Tariq Hassan (Platform) |