Satellite Storage Tiers

3-tier ABC storage, dedup, S3 presigned URLs, data lifecycle.

NIST controls:MP-4, SC-28, SI-12
Last updated:2026-03-22
Category:Technical
# Satellite Storage Architecture — Sensible Defaults by Scale ## Storage Backend Selection | Citus Mode | Default Storage | Why | Override | |------------|----------------|-----|----------| | **single** (1 node, ≤5K endpoints) | SeaweedFS | Simple, single-binary, low overhead, S3-compatible | Can choose CubeFS | | **ha** (3 nodes, ≤5K endpoints) | SeaweedFS (replicated) | SeaweedFS handles replication natively across volumes | Can choose CubeFS | | **multi-node** (3+ nodes, 5K+ endpoints) | **CubeFS** | Distributed POSIX + S3, scales horizontally with nodes, erasure coding | Can choose SeaweedFS | ## Why CubeFS for Multi-Node When a satellite scales to multi-node Citus: - **Citus workers need shared storage** for WAL archiving, backup, and data redistribution - **CubeFS** provides a distributed filesystem that grows with the cluster (add nodes → add capacity) - **Erasure coding** reduces storage overhead vs 3x replication (1.5x overhead for same durability) - **POSIX + S3 dual interface** — PostgreSQL WAL uses POSIX mount, agent uploads use S3 API - **CSI driver** — native K8s StorageClass, PVCs work transparently ## Why SeaweedFS for Single-Node - **Lower resource footprint** — master + volume + filer + s3 in ~500MB RAM total - **Simple operations** — single-node, no distributed consensus - **S3-compatible** — agent uploads work identically - **Already proven** — QA Test #1 validated the full pipeline ## CubeFS Configuration (Multi-Node Default) ```yaml # values-multi-node.yaml storage: backend: cubefs cubefs: enabled: true master: replicas: 3 # must match cluster node count storage: 10Gi metanode: replicas: 3 storage: 20Gi datanode: replicas: 3 storage: 100Gi # bulk block storage erasureCoding: true # 1.5x overhead vs 3x replication objectnode: replicas: 2 # S3 API gateway port: 8333 # same port as SeaweedFS for transparent swap csi: enabled: true storageClass: cubefs-sc seaweedfs: enabled: false ``` ## Migration Path: SeaweedFS → CubeFS When scaling from single-node to multi-node: 1. Deploy CubeFS alongside SeaweedFS (both running temporarily) 2. Migrate data: `aws s3 sync --endpoint-url <seaweedfs> s3://blocks/ --endpoint-url <cubefs> s3://blocks/` 3. Update S3 proxy to point to CubeFS objectnode 4. Verify agent uploads work against CubeFS 5. Decommission SeaweedFS 6. Total downtime: 0 (agent retries handle the cutover) ## Compliance Note Both SeaweedFS and CubeFS support encryption at rest (LUKS on the underlying volumes). For FedRAMP: - **SeaweedFS**: LUKS + volume-level encryption - **CubeFS**: LUKS + erasure coding (data fragments spread across nodes, no single node has complete data) CubeFS erasure coding provides additional security-through-distribution — a stolen disk contains only erasure-coded fragments, not complete files. ## Revision History | Date | Change | Author | |------|--------|--------| | 2026-03-22 | Initial — CubeFS default for multi-node, SeaweedFS for single | Daniel Okafor (Cloud Infrastructure), Tariq Hassan (Platform) |

This document is part of the Arivaran Twin compliance program. For questions or the latest version, contact compliance@arivaran.ai.

Release-ready. Saved on this browser.