A paper recovery design beside two real server racks and a stopwatch

RTO and RPO are business constraints expressed as engineering targets. They are not properties that an architecture diagram acquires because a document names them. A data platform meets an objective only when the complete recovery path produces measured evidence under the failure scenarios that matter.

The NIST Contingency Planning Guide defines the recovery point objective as the point in time to which data must be recovered after an outage. It defines the recovery time objective as the time an information system can remain in recovery before the outage harms the mission or business process. The definitions are short. Proving them is not.

Do not blend RTO and RPO into one promise

ObjectiveQuestion it answersEvidence I expectCommon false proof
RPOHow far back can recovered data be?Last durable recovery point, replication lag and data reconciliation resultA backup job completed successfully
RTOHow long can recovery take?Timed path from declared start to verified service restorationA standby server started
A platform can meet RTO and still fail RPO, or meet RPO and still fail RTO. Test and report them separately.

A replica that starts in five minutes may contain data older than the RPO. A backup that contains every committed record may take twelve hours to restore. Availability, recoverability and data currency overlap, but they are not the same metric.

Build a failure matrix before choosing the test

One successful restore does not validate every disaster. The test must match the failure boundary that the objective covers.

Failure scenarioRecovery path to testRPO evidenceRTO evidence
Single data node lossReplica or shard recovery while service stays onlineNo acknowledged data missingTime to restore redundancy and normal service margin
Availability zone lossTraffic shift, dependency failover and replica promotionLast durable record at the surviving siteDetection through verified application service
Network partitionWrite behavior on each side, fencing and reconciliationLag and conflict outcome during the partitionTime to safe service, not merely connectivity
Region or site lossAlternate-site activation and full dependency sequenceBackup or replica recovery pointEnd-to-end business service recovery
Control-plane lossCredential, DNS, orchestration and configuration recoveryConfiguration and secret recovery pointTime until operators can execute the data recovery path
The exact scenarios depend on the platform. The matrix forces the objective, mechanism and evidence to share the same failure boundary.

Chaos experiments and DR rehearsals answer different questions

Chaos engineering does not have to mean random failure injection. A good experiment is hypothesis driven and can test a precise partition, dependency loss or node failure. It is useful for checking whether the running system preserves an expected steady state.

A disaster recovery rehearsal has a different finish line. It measures the complete path from the agreed clock start to a verified business service, including decisions, credentials, DNS, certificates, dependency ordering, restore throughput, data reconciliation and client acceptance tests. I use both. I do not treat one as proof of the other.

Measure the complete recovery clock

  • Declare the start event. Hardware failure, monitoring alert, incident declaration and recovery authorization are different timestamps.
  • Include human decisions. Escalation, access approval, vendor support and failover authority consume real time.
  • Measure real throughput. Restore GB per hour from the production backup path, including decryption, decompression and validation.
  • Record dependency order. Identity, secrets, networking, DNS, certificates, databases, brokers and applications rarely recover in one step.
  • Validate data. A process exit code is not proof that the recovered dataset is complete and usable.
  • End at service verification. The clock stops when representative clients can complete critical operations and monitoring confirms stable behavior.
Claimed recovery timeline compared with a measured path containing delays and rework
The documented path usually omits waiting, failed prerequisites and validation. Only the measured path can support an RTO verdict.

Keep an evidence packet, not a pass or fail label

A useful rehearsal leaves enough evidence for another engineer to reproduce the verdict.

  • Scenario, assumptions, exclusions and target RTO and RPO.
  • Clock definition and timestamped event log.
  • Backup age, replication lag and restore throughput.
  • Commands, automation output and configuration versions.
  • Data validation queries and client acceptance results.
  • Observed dependency failures, manual work and vendor delays.
  • Measured RTO, measured recovery point and remaining uncertainty.
  • Corrective actions with owners and a retest date.
A framed recovery-plan diagram on the wall above a rack whose cables lie unplugged on the floor
The recovery plan assumes the cables; test the physical dependency order, not the diagram.

NIST places testing, training and exercises inside the contingency planning process and calls for scheduled validation as systems change. That matters for data platforms because a recovery result expires. Dataset growth changes restore time. New dependencies change ordering. Credential rotation can invalidate a runbook that worked six months ago.

Why independent validation matters

The implementation team should help design and run the rehearsal. It should not be the only party defining the pass criteria and interpreting ambiguous results. An independent review can test whether the target follows from the business impact analysis, whether the chosen scenario covers the claimed boundary and whether the evidence supports the conclusion.

Questions for the next design review

  • What exact event starts the RTO clock?
  • Which write defines the measured recovery point?
  • When was the last full restore from the real backup path?
  • How does replication behave during a partition or extended peer outage?
  • Which dependencies must recover first, and who owns each step?
  • What client operation proves that service is restored?
  • What evidence would make the reviewer reject the current RTO or RPO claim?

If the answers exist only in the architecture report, treat the objectives as unvalidated. I review recovery designs, runbooks, measurements and rehearsal evidence through a Resilience and Disaster Recovery Assessment. Book a 15-min intro call to discuss the platform and the failure boundary you need to prove.

Primary guidance

Update 2026-09-29. This post is now part of Data Platform Resilience: RTO, RPO and Replication in Practice, the field guide to RTO, RPO and replication.

Related posts


0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *