Backup And Restore Runbook
Backup Scope
Section titled “Backup Scope”Back up all of the following as one documented recovery set:
- Control PostgreSQL database, including
schema_migrations, identities, grants, source roles, policies, approvals, ingestion jobs, discovery metadata, cache dependencies, and audit partitions. - PostgreSQL WAL or provider PITR logs.
- Redis persistence when Admin session continuity and cache-generation state are in scope.
- Every Gateway audit spool persistent volume.
- External secret-store versions and encryption-key references.
- Helm values with secret values removed, chart version, image digest, and release manifest.
- Persisted vector/index data or the records and model metadata required to rebuild it.
Never place unencrypted database dumps, private keys, or source credentials in the repository or certification reports.
Schedule
Section titled “Schedule”Reference minimum:
- PostgreSQL PITR continuously, full backup daily, retained per policy.
- Redis persistence continuously and provider snapshot at least daily.
- Audit spool monitored continuously and included in volume snapshots at least hourly until delivered.
- Release manifests and SBOM retained for every deployed version.
- Restore rehearsal quarterly and before V1 promotion.
The actual interval must satisfy the selected RPO.
Pre-Backup Checks
Section titled “Pre-Backup Checks”- Confirm migration head and database health.
- Confirm audit partitions cover current and future months.
- Confirm audit queue/spool backlog and filesystem usage.
- Record release SHA, image digest, chart version, and backup timestamp.
- Confirm encryption and destination retention policy.
Isolated Restore Procedure
Section titled “Isolated Restore Procedure”- Create a network-isolated recovery environment with no routes to live data sources.
- Restore the control PostgreSQL backup to a new database instance.
- Restore Redis only when its snapshot is known compatible; otherwise start empty and expect session/cache loss.
- Attach copies of audit spool volumes. Never mount a production spool read-write into a rehearsal.
- Deploy the exact image digest that produced the backup.
- Verify
schema_migrationschecksums before applying newer migrations. - Replay audit spool events and verify idempotent event IDs prevent duplicate logical audit records.
- Reload or rebuild discovery vectors using matching embedding dimensions.
- Start services with external source egress disabled.
- Verify
/ready, Admin login, role/policy reads, audit history, approval state, worker queue state, and discovery metadata.
Recovery Validation
Section titled “Recovery Validation”The restore is not complete until:
- Row counts and critical-table checksums match the selected recovery point.
- No migration checksum mismatch exists.
- Active grants and policies produce expected allow and deny decisions.
- Audit events through the recovery point are queryable.
- Spool replay leaves no unexplained failed/dead-letter event.
- Pending approvals do not execute during validation.
- Secrets were restored by reference and no plaintext secret appears in logs.
- Measured RPO/RTO is recorded.
Production Recovery
Section titled “Production Recovery”After incident command approves recovery:
- Freeze writes and connector sync where possible.
- Select the recovery point based on integrity, not merely recency.
- Restore into replacement services rather than overwriting the only copy.
- Run isolated validation.
- Rotate credentials if compromise is possible.
- Switch traffic using a controlled DNS/service change.
- Watch audit, cache generation, worker leases, errors, and source circuits.
- Preserve old volumes and logs until incident closure.
Failure And Escalation
Section titled “Failure And Escalation”If backup decryption, migration verification, audit replay, or integrity checks fail, stop. Do not route production traffic. Escalate under incident-response.md and retain all evidence.