Change Management Policy

Change control board, approval workflow, rollback procedures.

NIST controls:CM-1 through CM-9
Last updated:2026-03-15
Category:Governance
# Change Management Policy **Document ID**: KD-POL-007 **Owner**: Chief Technology Officer **Approved By**: CEO, Arivaran **Effective Date**: 2026-03-19 **Next Review Date**: 2027-03-19 (recomputed 2026-10-02 on the stated semi-annual cycle from the Effective Date; the 2026-09-19 occurrence was not recorded, and this date does not assert that a review has taken place) **Review Cycle**: Semi-annual **Classification**: Internal ## 1. Purpose This policy defines how changes to KuiperDesk infrastructure and application code are proposed, reviewed, tested, approved, and deployed. KuiperDesk is managed entirely through infrastructure-as-code (Ansible, OpenTofu, Helm) and Git-based workflows. The change management process leverages these tools to make changes auditable, reversible, and repeatable. The goal is not to slow down deployment but to ensure every change is traceable and that untested changes never reach production. ## 2. Scope This policy covers all changes to: - KuiperDesk application code (aaagent, aaHashSvc, aaTwinSvc, aaDash) - Infrastructure configuration (Ansible playbooks, OpenTofu modules, Helm charts) - Kubernetes manifests and cluster configuration - DNS records (Cloudflare via OpenTofu) - CI/CD pipeline configuration (Forgejo Actions workflows) - Container images deployed to production - Firewall rules, NetworkPolicy, and network overlay configuration - PKI configuration (cert-manager issuers, OpenBao policies) Out of scope: Customer-side aaagent configuration changes (managed by tenant administrators through aaDash). ## 3. Change Classification | Type | Definition | Approval | Examples | |------|-----------|----------|---------| | **Standard** | Pre-approved, low-risk, routine changes that follow an established procedure | Automated (CI/CD pipeline passes) | Dependency version bump, Helm chart value change within tested parameters, certificate rotation | | **Normal** | Changes that modify system behavior, add features, or alter security controls | PR review + CTO approval | New service deployment, RBAC policy change, NetworkPolicy modification, new Ansible role | | **Emergency** | Changes required to resolve a P1/P2 incident where the normal process would cause unacceptable delay | CTO verbal approval, retroactive PR within 24 hours | Hotfix for active security vulnerability, network containment during incident | ## 4. Policy Statements ### 4.1 Git-Based Change Control 4.1.1. All infrastructure and application changes are tracked in the Forgejo Git repository (git.arivaran.ai/arivaran/arivaran-ai). The Git commit history is the authoritative change log. 4.1.2. The `main` branch has branch protection enabled: - Direct pushes to `main` are prohibited - All changes enter `main` via pull request (PR) - PRs require at least one approving review - PRs must pass all CI checks before merge - Force pushes to `main` are prohibited 4.1.3. Developers work on feature branches named with the convention: `<type>/<short-description>` (e.g., `feat/cnpg-backup-cronjob`, `fix/networkpolicy-egress`, `chore/helm-chart-bump`). ### 4.2 CI/CD Pipeline (Forgejo Actions) 4.2.1. Every PR triggers the following CI pipeline stages: **Stage 1 - Lint and Scan**: - `gitleaks`: Scans for hardcoded secrets, API keys, and credentials in the diff - `trivy`: Scans container images referenced in the change for known vulnerabilities - `ansible-lint`: Validates Ansible playbook syntax and best practices - Cargo clippy / npm lint for application code changes **Stage 2 - AI Review**: - Claude API review via OAuth token (CLAUDE_OAUTH_REFRESH_TOKEN) analyzes the change for security implications, configuration errors, and compliance impact **Stage 3 - Plan**: - `tofu plan` for OpenTofu changes (shows what will be created/modified/destroyed) - `helm diff` for Helm chart changes (shows manifest differences) - `ansible-playbook --check --diff` for Ansible changes (dry run) **Stage 4 - Deploy** (only after merge to main): - `tofu apply` for infrastructure changes - `helm upgrade` for Kubernetes workloads - `ansible-playbook` for OS-level configuration 4.2.2. If any Stage 1 check fails, the PR cannot be merged. There are no overrides for security scan failures without a documented exception (see Section 7). 4.2.3. The CI/CD runner executes on the control VM (10.0.0.40) using a Docker executor with the `ci-runner:v1` image from Harbor. The runner has network access to production systems for deployment stages only after merge. ### 4.3 Container Image Management 4.3.1. All container images deployed to production are stored in Harbor (harbor.arivaran.ai). Images pulled from upstream registries (Docker Hub, ghcr.io) are mirrored through Harbor's proxy cache. 4.3.2. Container images are signed using cosign. The CI pipeline signs images after successful build and scan. The Kubernetes admission controller (or runtime policy) verifies signatures before allowing pod creation. 4.3.3. Helm chart versions are pinned in the deployment configuration. Upgrades to chart versions are treated as Normal changes requiring PR review. 4.3.4. Base image versions are pinned to specific digests, not floating tags. Upstream base image updates are tracked and applied through the normal change process. ### 4.4 Change Review Requirements 4.4.1. PR reviews must verify: - The change achieves its stated objective - No unintended side effects on other services or tenants - Security controls are not weakened (RBAC, NetworkPolicy, encryption, authentication) - Rollback procedure is documented for Normal changes - Monitoring/alerting is updated if the change introduces new failure modes 4.4.2. Changes that modify security controls (RBAC, NetworkPolicy, firewall rules, PKI configuration, authentication flows) require explicit CTO approval in the PR review, documented with a comment referencing the relevant policy section. 4.4.3. Changes that modify data stores (PgEdge/CNPG schema, SeaweedFS configuration) require a data migration plan and rollback procedure documented in the PR description. ### 4.5 Deployment and Rollback 4.5.1. Deployments to production occur only from the `main` branch via the CI/CD pipeline. Manual kubectl/helm commands against production are prohibited except during emergency changes. 4.5.2. Helm deployments use `--atomic` flag: if the deployment fails health checks, it automatically rolls back to the previous release. 4.5.3. For Ansible changes, the `--check --diff` dry run output is reviewed before applying. Ansible playbooks are idempotent; re-running a previous version effectively rolls back. 4.5.4. Rollback procedure for Normal changes: revert the merge commit on `main`, which triggers the CI/CD pipeline to deploy the previous state. This is documented in the PR that introduced the change. ### 4.6 Emergency Changes 4.6.1. Emergency changes bypass the PR review requirement but must still be committed to Git. The procedure: 1. CTO verbally authorizes the change (documented in the incident Forgejo issue) 2. Engineer makes the change directly (kubectl, helm, ansible) and commits to a branch 3. Within 24 hours: a retroactive PR is created documenting what was changed, why, and the incident reference 4. The retroactive PR undergoes normal review 4.6.2. Emergency changes are tracked and reported at the quarterly risk review. A pattern of frequent emergency changes indicates a process or architecture problem requiring remediation. ### 4.7 Change Audit Trail 4.7.1. The following artifacts constitute the change audit trail: - Git commit history with author, timestamp, and diff - PR description, review comments, and approval record - CI/CD pipeline logs (lint, scan, plan, deploy stages) - Helm release history (`helm history`) - Ansible execution logs 4.7.2. These artifacts are retained for a minimum of 12 months and are available as SOC 2 evidence for the CC8.1 control objective. ## 5. Roles and Responsibilities | Role | Responsibility | |------|---------------| | CTO | Approve Normal changes to security controls. Authorize Emergency changes. Review change metrics quarterly. | | Infrastructure Engineers | Create PRs. Review peers' PRs. Execute deployments via CI/CD. Document rollback procedures. | | CI/CD Pipeline | Automated enforcement of lint, scan, plan gates. Blocks non-compliant changes. | ## 6. Compliance Mapping | Policy Statement | SOC 2 Criteria | ISO 27001:2022 | |-----------------|----------------|----------------| | 4.1 Git-Based Change Control | CC8.1 | A.8.9 | | 4.2 CI/CD Pipeline | CC8.1 | A.8.9, A.8.25 | | 4.3 Container Image Management | CC8.1 | A.8.9 | | 4.4 Change Review | CC8.1 | A.8.9 | | 4.5 Deployment and Rollback | CC8.1 | A.8.9 | | 4.6 Emergency Changes | CC8.1 | A.8.9 | | 4.7 Audit Trail | CC8.1, CC7.1 | A.8.9, A.8.15 | ## 7. Exceptions Exceptions to CI scan requirements (e.g., deploying with a known non-critical Trivy finding) require a Forgejo issue documenting: the finding, risk assessment, compensating controls, and CTO approval. The exception is time-bounded (maximum 30 days) and tracked to remediation. ## 8. Enforcement Changes deployed outside the CI/CD pipeline (except authorized Emergency changes) are policy violations. Drift detection compares live cluster state against Git-declared state weekly; any drift is investigated and corrected. ## 9. Revision History | Version | Date | Author | Changes | |---------|------|--------|---------| | 1.0 | 2026-03-19 | CTO | Initial release |

This document is part of the Arivaran Twin compliance program. For questions or the latest version, contact compliance@arivaran.ai.

Release-ready. Saved on this browser.