Upgrade And Rollback Runbook
Preconditions
Section titled “Preconditions”- Target image/chart is signed, attested, vulnerability-scanned, and identified by immutable digest.
- Release notes identify schema and compatibility changes.
- Current release is within the supported upgrade window.
- A restorable pre-upgrade PostgreSQL backup and spool snapshot exist.
- Migration was tested against a restored production-shaped database.
- Capacity, maintenance window, rollback owner, and success criteria are set.
Upgrade Procedure
Section titled “Upgrade Procedure”- Record current release SHA, image digest, chart version, migration head, and health status.
- Pause nonessential ingestion and resolve or record long-running approvals.
- Confirm audit queue/spool is healthy and below capacity thresholds.
- Take/verify the pre-upgrade backup.
- Render the target Helm chart with production values and policy checks.
- Run the pre-upgrade migration Job. The advisory lock prevents concurrent migration runners, but not conflicting operator actions.
- Verify migration checksums and head.
- Roll out Gateway, Admin, and Workers using the same image digest.
- Wait for dependency-aware readiness before admitting traffic.
- Run smoke paths: authentication, allow, deny, redaction, cache miss/hit, queued approval, rejection, audit detail, and discovery.
- Resume ingestion and monitor for one observation window.
Configuration-Only Changes
Section titled “Configuration-Only Changes”A helm upgrade that changes only ConfigMap-backed values rolls the Gateway,
Admin, and Workers through the checksum/config pod annotation; confirm the
rollout completes and readiness recovers as in steps 8-9.
Rotating a Secret, or changing anything supplied through extraEnvFrom, does
not roll anything, because the chart cannot see those objects. After such a
change run kubectl rollout restart for each affected workload and verify the
running environment picked it up.
Upgrading From A Release Before 1.0.0-rc.5
Section titled “Upgrading From A Release Before 1.0.0-rc.5”Charts before 1.0.0-rc.5 put per-release labels (helm.sh/chart,
app.kubernetes.io/version) on the Gateway StatefulSet’s volumeClaimTemplates.
Kubernetes forbids changing that field, so upgrading such a deployment to any
newer chart fails and Helm rolls back:
StatefulSet.apps "<release>-gateway" is invalid: spec: Forbidden: updates tostatefulset spec for fields other than 'replicas', ... are forbiddendeploy/scripts/release/deploy-oci-release.sh handles the transition. Before
helm upgrade it reads the live Gateway claim template, and only if it still
carries one of those labels it runs:
kubectl --namespace <namespace> delete statefulset <release>-gateway --cascade=orphanThis deletes the StatefulSet object alone. The Gateway pod and its audit spool
volume keep running and serving, and the StatefulSet helm upgrade recreates
adopts them because its selector is unchanged. The volume keeps its data. Run
the same command yourself before helm upgrade if you do not use the script.
It happens once: charts from 1.0.0-rc.5 onward carry only stable labels there.
Rollback Decision
Section titled “Rollback Decision”Rollback when any of the following persists beyond the defined observation threshold:
- Mandatory readiness cannot become healthy.
- Authorization, redaction, write safety, or audit is incorrect.
- Error/latency/resource rates exceed the release threshold.
- Worker leases or audit spool backlog grow without recovery.
- Data integrity or migration verification is uncertain.
Security/integrity defects take precedence over availability. Stop governed traffic if continuing could expose or corrupt data.
Application-Only Rollback
Section titled “Application-Only Rollback”Use Helm/application rollback only if release notes state the previous version is compatible with the new schema.
- Freeze new writes.
- Roll back all runtime roles to the prior immutable digest.
- Verify readiness and migration compatibility.
- Run the smoke matrix.
- Resume traffic and retain incident evidence.
Database Restore Rollback
Section titled “Database Restore Rollback”If the prior application cannot read the migrated schema:
- Stop Gateway, Admin mutations, Workers, and migration jobs.
- Preserve the failed database and audit spool for investigation.
- Restore the verified pre-upgrade database to a replacement instance.
- Reconcile audit spool events created after the backup. Do not blindly replay writes or approvals.
- Deploy the prior application digest.
- Validate grants, policies, approval terminal states, audit continuity, and cache generations before traffic.
V1 has no automatic down-migration contract.
Post-Change Record
Section titled “Post-Change Record”Record timestamps, operators, backup ID by safe reference, old/new digests, migrations applied, smoke results, rollback decision, and unresolved issues. Do not record secrets, raw tokens, or private payloads.