A Disaster Recovery (DR) test should answer a defined question about an important service. Restoring a selection of virtual machines, completing a storage failover or confirming that backup jobs are healthy may be useful technical checks, but none of them necessarily shows that users can regain a working service.
A proper test connects the disruption scenario, recovery target, service dependencies, people, procedures and acceptance checks. It records what happened, explains what remained outside scope and produces actions that are retested.
Start with the recovery question
“Test DR” is not a sufficient objective. The team needs to agree what event it is preparing for and what the exercise must demonstrate.
Examples include:
- recovering a customer-facing service after the loss of its primary site;
- restoring an application and its data after destructive change;
- operating from a secondary platform while a primary environment is unavailable;
- recovering administrative access when normal identity services cannot be used; or
- confirming that a critical supplier dependency has a workable continuity route.
The scenario determines which controls, people and dependencies matter. A storage failure, ransomware event and regional cloud disruption place different demands on the same service.
State the assumed loss clearly. If the test allows access to systems, credentials or documentation that would be unavailable during the scenario, record that as a limitation rather than quietly simplifying the exercise.
Define the service, not only the infrastructure
Infrastructure scope often begins with servers, virtual machines, databases or storage volumes. Service scope asks what must work together before the organisation can use the service again.
That may include:
- identity, privileged access and recovery credentials;
- DNS, routing, firewalls and load balancing;
- certificates, secrets and service accounts;
- application components and data in the required sequence;
- third-party connections and external dependencies;
- monitoring, logging and backup services; and
- people who can authorise recovery and confirm service acceptance.
The recovery scope should name the business and technical owners, the components included, explicit exclusions and any shared foundation being assumed. If one missing dependency would prevent service use, it belongs in the recovery discussion even when another team owns it.
Set observable success criteria
A useful success criterion can be tested and evidenced. “The application came back” leaves too much open to interpretation.
Better criteria might state that:
- the agreed service functions were available to an authorised test user;
- data was recovered to a stated point and reconciled by the responsible owner;
- authentication, name resolution and external connections worked through the recovery route;
- monitoring and operational alerting were active;
- recovery completed within the measured target; and
- the business or service owner accepted the result.
Success criteria should also identify who decides whether the test has passed. Infrastructure health alone cannot confirm that a business service is usable.
Choose the right depth of test
Not every exercise needs a full production failover. The appropriate test depends on the question, risk and available controls.
A technical restore test can demonstrate that selected data or components can be retrieved. It is valuable, but its conclusion should remain limited to those components.
A service recovery exercise brings together the service components and material dependencies in a representative recovery route. It should include application and service validation.
A procedural or tabletop exercise tests decisions, responsibilities, escalation and the usability of plans without operating the recovery technology. It can expose serious gaps but cannot demonstrate technical recovery time.
A live failover exercise can provide stronger operational evidence, but it requires suitable change controls, rollback arrangements, monitoring and authority. It should not be treated as the only respectable form of testing.
The depth should be stated before the exercise so that a limited test is not later reported as broader assurance.
Prepare the recovery route and safety controls
The exercise plan should identify prerequisites, expected sequence, decision points, safety boundaries and rollback conditions. This is not about scripting every action so tightly that the test becomes artificial. It is about preventing avoidable production risk and making the result interpretable.
Confirm before the exercise:
- who can start, pause and stop the test;
- which production changes are authorised;
- which credentials, tools and communication channels are available;
- how the team will distinguish test activity from an incident;
- what will be monitored during recovery; and
- how the environment will be returned to its normal state.
The people expected to recover the service should be involved. If the procedure only works when its original author is present, the exercise has identified an operational dependency.
Measure the whole recovery sequence
Record when the scenario begins, when recovery is authorised, when technical work starts, when each dependency becomes available and when the service is accepted. These timestamps distinguish hands-on restore time from total service recovery time.
Do not remove delays simply because they are organisational rather than technical. Locating the right person, obtaining emergency access, approving a network change or waiting for a supplier can determine the real outcome.
The companion guide Your RTO says four hours. Can you prove it? explains how to compare the stated recovery time objective with the measured sequence.
Capture evidence while the test is running
Screenshots assembled days later are rarely enough. Capture evidence against the planned steps and success criteria while the exercise is in progress.
Useful records include:
- start, decision and completion timestamps;
- backup, restore, replication and failover output;
- relevant configuration and health records;
- application and user-validation results;
- exceptions, workarounds and unexpected dependencies;
- decisions made and the people authorised to make them; and
- failed steps, repeated actions and unresolved questions.
The evidence should make it possible for somebody outside the exercise team to understand what was attempted and the basis for the conclusion. What evidence should a Disaster Recovery test produce? sets out a practical evidence pack.
State what the result does and does not demonstrate
A test conclusion should be bounded by the service, scenario, recovery route and evidence available. If authentication was bypassed, a dependency was pre-started or data validation was simulated, say so directly.
This does not make the exercise a failure. It prevents a useful partial result from being overstated.
A concise conclusion might say that the selected application components were restored successfully within the technical target, while end-to-end service recovery remains unproven because external authentication and business validation were outside scope.
That is more useful than a green status that conceals the remaining question.
Turn findings into owned actions and retest
The exercise is not complete when the service returns. Findings need an owner, priority, target date and evidence required for closure.
Separate actions that affect recovery capability from general improvement. A missing emergency credential, unusable runbook or untested data-reconciliation step should not disappear into a long operational backlog.
Retest changes that materially affect the recovery conclusion. Updating a document or closing a ticket does not by itself show that the improved route works.
A practical starting point
Choose one important service and write down the disruption scenario, success criteria, components, dependencies and evidence you would expect a representative exercise to produce. Compare that expectation with the most recent test record.
The DR Readiness Checklist provides a structured set of prompts. The DR Assurance Sample Report shows how exercise evidence, limitations and actions can be presented, while Technical Disaster Recovery Assurance explains the independent review scope and indicative fees.