# Incident Response Policy
**Document ID**: KD-POL-003
**Owner**: Chief Technology Officer
**Approved By**: CEO, Arivaran
**Effective Date**: 2026-03-28
**Next Review Date**: 2027-03-28 (recomputed 2026-10-02 on the stated semi-annual cycle from the Effective Date; the 2026-09-28 occurrence was not recorded, and this date does not assert that a review has taken place)
**Review Cycle**: Semi-annual
**Classification**: Internal
## 1. Purpose
This policy defines how Arivaran detects, responds to, and recovers from security incidents affecting KuiperDesk. A backup-as-a-service platform is a high-value target: compromised backup data can be used for extortion, and compromised restore capabilities can destroy a customer's disaster recovery plan. The incident response process must be fast, structured, and auditable.
## 2. Scope
This policy covers security incidents affecting:
- KuiperDesk infrastructure (RKE2 cluster, VPS nodes, network overlay, CI/CD pipeline)
- Customer backup data stored in S3-compatible storage (SeaweedFS), metadata in PostgreSQL (PgEdge on hub, CNPG on satellites), catalog data in PostgreSQL/Citus
- Agent-to-service communication channels
- Identity and access management systems (Keycloak, OpenBao)
- Supply chain incidents affecting KuiperDesk dependencies
An incident is any event that compromises or threatens the confidentiality, integrity, or availability of KuiperDesk systems or customer data. Near-misses and policy violations are also handled through this process.
## 3. Incident Severity Classification
| Severity | Definition | Response SLA | Example |
|----------|-----------|-------------|---------|
| **P1 - Critical** | Active data breach, service-wide outage, or ransomware. Customer data confirmed or likely compromised. | Acknowledge: 15 min. Contain: 1 hour. | Unauthorized access to PostgreSQL (PgEdge/CNPG)/SeaweedFS, RKE2 node compromise, supply chain attack on deployed image, WORM bucket lock disabled |
| **P2 - High** | Attempted breach blocked but attack ongoing, single-tenant impact, degraded security controls. | Acknowledge: 30 min. Contain: 4 hours. | Brute force against Keycloak with partial success, single node offline, certificate compromise for one agent |
| **P3 - Medium** | Policy violation, vulnerability discovered in production, failed security control. | Acknowledge: 4 hours. Remediate: 48 hours. | Unpatched CVE in running container, misconfigured NetworkPolicy, failed access review finding |
| **P4 - Low** | Security anomaly under investigation, informational finding, minor policy deviation. | Acknowledge: 24 hours. Remediate: 1 week. | Unusual but authorized access pattern, non-critical audit log gap |
## 4. Policy Statements
### 4.1 Detection
4.1.1. **Runtime Security (Tetragon)**: Tetragon is deployed as a DaemonSet on all RKE2 nodes. Tracing policies detect:
- Unexpected process execution in container namespaces (e.g., shell spawned in aaHashSvc pod)
- File access to sensitive paths (/etc/shadow, /var/lib/rancher/rke2/server/token, /run/secrets)
- Network connections to unexpected destinations from workload pods
- Privilege escalation attempts (setuid, capability changes)
Tetragon events are forwarded to Loki and trigger AlertManager rules for P1/P2 patterns.
4.1.2. **Log-Based Detection (Loki + AlertManager)**: AlertManager rules are configured for:
- Authentication anomalies: 5+ failed logins in 10 minutes, login from an IP outside the approved admin / satgw-tunnel range, break-glass account use
- Infrastructure anomalies: Node NotReady, pod crash loops in security-critical workloads, etcd leader changes
- Certificate anomalies: Unexpected certificate issuance, CRL update outside of scheduled operations
- Data anomalies: aaHashSvc receiving requests for hash ranges outside a tenant's scope, bulk deletion requests exceeding threshold
4.1.3. **External Monitoring (Gatus)**: The status page at status.arivaran.ai monitors endpoint availability for all 8 public services. Downtime triggers PagerDuty alerts (when integrated) and is visible to customers.
4.1.4. **Vulnerability Scanning**: Trivy scans all container images in the CI/CD pipeline before deployment. Weekly scans of running images detect newly disclosed CVEs. Critical/High CVEs in production images generate P3 incidents automatically.
### 4.2 Incident Response Phases
#### Phase 1: Identification and Triage
4.2.1. The on-call engineer (currently the CTO given team size) receives the alert, confirms it is a genuine incident (not a false positive), and assigns a severity level per Section 3.
4.2.2. A Forgejo issue is created with the label `incident` and severity tag. The issue template captures: detection time, initial indicators, affected systems, initial severity assessment.
4.2.3. For P1/P2 incidents, the CTO is the Incident Commander. All actions and decisions are logged in the Forgejo issue in real time.
#### Phase 2: Containment
4.2.4. **Network Containment**: Apply Cilium NetworkPolicy to isolate compromised pods. For node-level compromise, remove the node from the RKE2 cluster (`kubectl drain` + firewalld block). Revoking the node's satgw-tunnel credential isolates it from cross-node services.
4.2.5. **Identity Containment**: Disable compromised Keycloak accounts immediately. For agent certificate compromise, add the certificate serial to the CRL via OpenBao and trigger a CRL distribution update. aaHashSvc and aaTwinSvc validate the CRL on each TLS handshake.
4.2.6. **Data Containment**: If unauthorized data access is confirmed, snapshot the affected PostgreSQL cluster (PgEdge on hub, CNPG on satellites) and SeaweedFS volumes for forensic analysis before any remediation that might alter evidence.
4.2.7. Short-term containment prioritizes stopping ongoing damage. Long-term containment may involve deploying a patched version, rotating credentials, or re-issuing certificates, which happens in parallel with investigation.
#### Phase 3: Eradication
4.2.8. Identify and remove the root cause. For compromised containers: rebuild from verified base images and redeploy. For compromised nodes: reimage from Ansible playbooks (infrastructure-as-code ensures known-good state). For compromised credentials: rotate all credentials that may have been exposed, not just confirmed ones.
4.2.9. Verify eradication by checking Tetragon traces and Loki logs confirm no further malicious activity for a minimum observation period of 24 hours (P1) or 4 hours (P2).
#### Phase 4: Recovery
4.2.10. Restore services in order of priority: (1) Keycloak (identity), (2) OpenBao (certificates), (3) PostgreSQL (PgEdge on hub, CNPG on satellites), (4) aaHashSvc + aaTwinSvc (application), (5) aaDash (dashboard).
4.2.11. Verify data integrity using content-addressed block verification (SHA-256 hash comparison) for any backup data that may have been modified.
4.2.12. Monitor recovered systems with increased alerting sensitivity for 7 days post-recovery.
#### Phase 5: Post-Mortem
4.2.13. A blameless post-mortem is required for all P1 and P2 incidents and recommended for P3. The post-mortem is conducted within 5 business days of incident closure and documented in the Forgejo issue.
4.2.14. Post-mortem template:
- Timeline of events (detection through resolution, with timestamps)
- Root cause analysis (5 Whys or equivalent)
- What went well in the response
- What could be improved
- Action items with owners and deadlines
4.2.15. Action items from post-mortems are tracked in Forgejo and reviewed at the next quarterly risk meeting. Systemic issues result in policy or control updates.
### 4.3 Communication and Notification
4.3.1. **Internal Communication**: P1/P2 incidents are communicated to all Arivaran personnel within 1 hour of confirmation. Communication channels: Jitsi for real-time coordination, Forgejo issue for persistent record.
4.3.2. **Customer Notification**: Affected customers are notified within 24 hours of confirming a P1 incident that impacts their data or service availability. Notification includes: what happened, what data was affected, what actions Arivaran is taking, and what actions the customer should take.
4.3.3. **Regulatory Notification (GDPR)**: If a personal data breach is confirmed, the relevant supervisory authority is notified within 72 hours per GDPR Article 33. The notification is prepared by the CTO and reviewed by legal counsel. Data subjects are notified without undue delay if the breach is likely to result in high risk to their rights (Article 34).
4.3.4. **Auditor Notification**: Material incidents are disclosed to the SOC 2 auditor as part of the next audit cycle. Incidents that result in a control failure during the audit period are proactively reported.
### 4.4 Data Retention and WORM Violation Incidents
WORM (Write Once, Read Many) protection is a core security control for KuiperDesk backup data. Violations or anomalies in the WORM subsystem are treated as security incidents because they may indicate tampering, ransomware activity, or infrastructure misconfiguration that undermines data immutability guarantees.
#### 4.4.1 WORM Incident Types and Severity
| Incident Type | Trigger | Severity | Rationale |
|---------------|---------|----------|-----------|
| **Bucket Object Lock disabled** | `WormBucketLockDisabled` alert fires (`kuiperdesk.hashsvc.worm.bucket_lock_enabled == 0`) | **P1 - Critical** | All WORM protection is ineffective. Every block uploaded without Object Lock can be deleted by any privileged user. WormWorker refuses to start. |
| **Persistent unprotected blocks** | `WormBlocksUnprotected` alert fires (`blocks_unprotected > 0` for 30+ min) | **P1 - Critical** | Blocks without S3 Object Lock are vulnerable to deletion. If WormWorker cannot converge to zero, there is a systemic failure in lock application. |
| **WormWorker prolonged failure** | `WormWorkerDown` alert fires (metric absent for 20+ min) or WormWorker fails to start | **P1 - Critical** | Without WormWorker, new blocks are not locked and existing locks are not extended. A silent failure here creates a growing window of unprotected data. |
| **Elevated delete-blocked events** | `WormDeleteBlocked` alert fires (>5 blocked deletions/hr per tenant) | **P2 - High** | May indicate compromised admin credentials attempting to delete backups (ransomware pattern), or a misconfigured automation repeatedly attempting illegal deletions. |
| **S3 lock extension failures** | `WormLockExtensionFailures` alert fires (non-zero failure rate for 15+ min) | **P2 - High** | S3 putObjectRetention calls are failing. If the current lock expires before extension succeeds, blocks become deletable. May indicate S3 endpoint issues, IAM policy changes, or network partition. |
| **S3/DB lock drift detected** | Periodic reconciliation finds blocks where S3 Object Lock retention differs from `hash_metadata.lock_until` by more than 24 hours | **P2 - High** | Indicates either a failed DB update after successful S3 call (transaction failure), or S3 Object Lock was modified outside the application (unauthorized access to S3 admin API). |
| **Audit hash chain break** | `verify_audit_chains()` (every service DB; aaTwinSvc's `verify_audit_hash_chain()` delegates to it) returns a `HASH_MISMATCH`, `SEQ_GAP`, `BROKEN_LINK`, `DUPLICATE_SEQ` or `MISSING_GENESIS` row | **P1 - Critical** | Tamper evidence: a row inside one writer's chain was rewritten, deleted or reordered. This may indicate a compromised database superuser or storage-level tampering. Since #22538 each writer has its own chain, so these are no longer produced by replicas interleaving; rows written before the cut-over (`chain_id IS NULL`) are counted as legacy and never reported as breaks. |
| **Retention shortening attempt** | `trg_snapshot_worm` trigger fires and blocks an UPDATE reducing `retention_expires_at` | **P3 - Medium** | The database trigger prevented the shortening, so no data impact occurred. Investigate who or what issued the SQL UPDATE. May be a bug in application code or a manual DBA action. |
#### 4.4.2 WORM Incident Response Procedures
**For P1 WORM incidents (Bucket Lock disabled, persistent unprotected blocks, WormWorker down, audit chain break):**
1. **Immediate containment (within 15 minutes)**:
- Verify the alert is genuine (not a monitoring false positive)
- If bucket Object Lock is disabled: halt all new backup ingestion to prevent creating unprotected blocks. Investigate whether the bucket was recreated without Object Lock.
- If WormWorker is down: restart aaHashSvc pod. If WormWorker refuses to start (bucket Object Lock check failed), treat as bucket lock disabled.
- If audit chain is broken: isolate the PostgreSQL cluster (PgEdge/CNPG) from write access. Snapshot the database for forensic analysis before any remediation.
2. **Investigation**:
- Review aaHashSvc and aaTwinSvc logs for the 24 hours preceding the alert
- Check S3 bucket configuration change history (if available from S3 admin audit log)
- Check Keycloak and PostgreSQL (PgEdge/CNPG) audit logs for unauthorized access
- For unprotected blocks: query `SELECT COUNT(*) FROM hash_metadata WHERE lock_until IS NULL` and identify the tenant scope
3. **Remediation**:
- Bucket lock disabled: Object Lock can only be enabled at bucket creation time. If the bucket was recreated without it, data must be migrated to a new Object Lock-enabled bucket. This is a major incident requiring data migration planning.
- WormWorker failure: fix the root cause (S3 connectivity, IAM permissions, configuration) and restart. After restart, verify the unprotected block count converges to zero within one WormWorker cycle.
- Audit chain break: restore from the last verified backup of the audit partition. Cross-reference with the S3-archived audit copies (which are Object Lock protected and trustworthy).
4. **Post-incident**: Full post-mortem per Section 4.2.13-4.2.15.
**For P2 WORM incidents (elevated delete-blocked, lock extension failures, S3/DB drift):**
1. **Triage (within 30 minutes)**:
- For delete-blocked: identify the tenant and user issuing the deletion requests from aaTwinSvc audit logs. Determine if this is a compromised account or misconfigured automation.
- For lock extension failures: check S3 endpoint health, IAM policy, and network connectivity from the aaHashSvc pod.
- For S3/DB drift: run reconciliation query comparing `hash_metadata.lock_until` against actual S3 Object Lock retention (sample of blocks).
2. **Containment**:
- For suspected compromised credentials: disable the Keycloak account immediately and rotate the tenant's API tokens.
- For S3/DB drift: the S3 Object Lock is the authoritative source. Update `hash_metadata.lock_until` to match S3 reality. Investigate how the drift occurred.
3. **Resolution**: Fix root cause, verify metrics return to normal, document in Forgejo.
**For P3 WORM incidents (retention shortening blocked by trigger):**
1. Investigate the source of the SQL UPDATE that attempted to shorten retention
2. If from application code: file a bug, fix, and deploy
3. If from a DBA session: review with the DBA, document the reason, and remind of WORM policy
4. No data was impacted (the trigger prevented the change)
#### 4.4.3 Periodic WORM Reconciliation
4.4.3.1. A weekly CronJob shall perform S3/DB lock reconciliation by sampling 1% of blocks per tenant and comparing `hash_metadata.lock_until` against the actual S3 Object Lock `RetainUntilDate`. Discrepancies exceeding 24 hours are reported as P2 incidents.
4.4.3.2. The weekly `verify_audit_hash_chain()` function call is run by a Kubernetes CronJob. Any `BROKEN_LINK` result triggers a P1 incident automatically via AlertManager.
### 4.5 Tabletop Exercises
4.5.1. Tabletop incident response exercises are conducted semi-annually. Scenarios shall include at least: (a) compromised agent certificate used to exfiltrate backup data, (b) RKE2 node compromise via container escape, (c) supply chain attack on a container image, (d) ransomware actor with compromised admin credentials attempting to delete WORM-protected backups, (e) S3 bucket Object Lock found disabled during routine verification.
4.5.2. Exercise findings are documented and result in response procedure updates.
## 5. Roles and Responsibilities
| Role | Responsibility |
|------|---------------|
| CTO / Incident Commander | Declare incident severity. Coordinate response. Approve customer notifications. Authorize containment actions. |
| Infrastructure Engineers | Execute containment and eradication. Collect forensic evidence. Implement recovery procedures. |
| CEO | Approve regulatory notifications. Authorize external communications. |
## 6. Compliance Mapping
| Policy Statement | SOC 2 Criteria | ISO 27001:2022 | FedRAMP |
|-----------------|----------------|----------------|---------|
| 4.1 Detection | CC7.1, CC7.2 | A.8.15, A.8.16 | SI-4 |
| 4.2 Response Phases | CC7.3, CC7.4 | A.5.24, A.5.25, A.5.26 | IR-4, IR-5 |
| 4.3 Communication | CC7.4, CC7.5 | A.5.27, A.6.8 | IR-6 |
| 4.4 WORM Violation Incidents | CC7.2, CC7.3 | A.8.16, A.5.24 | IR-4, AU-6, SC-28 |
| 4.5 Tabletop Exercises | CC7.3 | A.5.28 | IR-3 |
## 7. Exceptions
There are no exceptions to the incident response process. All suspected incidents must be triaged. Severity classification may be adjusted during investigation, but the process itself is mandatory.
## 8. Enforcement
Failure to report a known or suspected security incident is a policy violation subject to disciplinary action. Deliberate concealment of an incident is grounds for immediate termination.
## 9. Revision History
| Version | Date | Author | Changes |
|---------|------|--------|---------|
| 1.0 | 2026-03-19 | CTO | Initial release |
| 1.1 | 2026-03-28 | CTO | Add Section 4.4 WORM violation incident types, severity classification, response procedures, and periodic reconciliation (FedRAMP IR-4). Update TiKV references to PostgreSQL (PgEdge on hub, CNPG on satellites). Add WORM tabletop exercise scenarios. Add FedRAMP column to compliance mapping. |