How to build a Disaster Recovery plan on AWS: RTO and RPO in practice
Building a disaster recovery plan is not just choosing a backup tool and trusting that it will work when you need it. It is a structured process, which starts by understanding how much downtime and data loss the operation can tolerate, and ends with a truly tested architecture. This guide shows the main steps for building this plan on AWS.
The starting point: RTO and RPO
Before choosing any tool, it is necessary to define two numbers for each critical application of the company.
RTO, or Recovery Time Objective, is the maximum acceptable time between the failure and the operation coming back online. If a system’s RTO is four hours, this means the company can tolerate up to four hours of downtime without serious impact on the business.
RPO, or Recovery Point Objective, is the maximum amount of data the company can lose, measured in time. An RPO of fifteen minutes means that, in the worst case, the company loses at most fifteen minutes of data since the last recovery point.
Not every application needs the same level of protection. An e-commerce system that generates direct revenue probably needs low RTO and RPO, measured in minutes. An internal reporting system may tolerate hours of downtime without major loss.
The four most used DR strategies on AWS
AWS organizes disaster recovery strategies into four categories, ordered from lowest to highest cost, and from highest to lowest RTO.
- Backup and restore is the simplest and cheapest strategy. The data is stored securely, usually with AWS Backup, and the infrastructure is only recreated when the disaster happens. The RTO is usually hours.
- Pilot light keeps a minimal, essential version of the infrastructure always active in a secondary region, such as the replicated database, while the rest of the resources are provisioned only when necessary. The RTO drops to tens of minutes.
- Warm standby keeps a reduced but functional version of the complete environment running all the time in a second region. In case of failure, the environment is scaled quickly to support the full traffic. The RTO is in the range of minutes.
- Multi-site active-active keeps the complete environment running simultaneously in two or more regions, distributing traffic among them all the time. It is the strategy with an RTO closest to zero, and also the most expensive to maintain.
AWS tools to put the plan into practice
AWS Backup centralizes and automates the backup policy of various services, such as EC2, RDS, EFS, and DynamoDB, in a single panel, with retention and cross-region copy rules. AWS Elastic Disaster Recovery, known as DRS, replicates entire servers, including those running outside AWS, to a target region, allowing failover in minutes without needing to keep a mirrored infrastructure running all the time. Amazon S3 with cross-region replication protects data stored in buckets against the unavailability of an entire region. Amazon Route 53 with health checks and automatic failover redirects traffic to the contingency environment as soon as it detects that the main environment is offline.
How to choose the right strategy?
The choice depends on three factors: how much the application is worth to the business, what RTO and RPO it really requires, and what budget is available to maintain this protection. Critical revenue applications usually justify warm standby or multi-site active-active. Internal support systems are usually well served by backup and restore or pilot light.
The plan is only worth it if it is tested
A disaster recovery plan that has never been tested in practice is an assumption, not a guarantee. Periodic failover simulations show whether the promised RTO is really achieved, and avoid surprises precisely at the moment when the company most needs everything to work.
CloudDog helps companies design, implement, and test disaster recovery strategies on AWS, with RTO and RPO defined according to the real criticality of each application. Learn about our Backup and Disaster Recovery service and talk to our architects to find out which strategy makes sense for your environment.

