Skip to content

Business Continuity

A recovery plan that has never been run is a document, not a capability.

What this is

Business continuity and disaster recovery is the work of deciding what must survive, how quickly it must return, and proving that it does. It spans impact analysis, RPO and RTO targets agreed with the business, failover design, runbooks and exercises.

The recurring failure is a plan built around infrastructure rather than around services. Everything is replicated, and at the exercise it emerges that the order of return was never established, the DNS change needs someone who has left, and the recovery environment cannot carry full production load. Recovery is a sequence, and untested sequences do not hold.

When you need it

If more than one of these is true, this is usually the right place to start.

  • A DR plan exists, was written some years ago, and has never been exercised end to end.
  • RPO and RTO numbers sit in a document that no business owner ever agreed to.
  • A regulator, insurer or customer has asked for evidence of tested recovery and there is none to give.
  • An incident or near miss revealed that the recovery steps live in one person's head.

What the scope covers

  • Business impact analysis: which services matter, what an hour of loss costs, and the dependencies each one carries.
  • RPO and RTO targets per service, agreed with the business owners who bear the consequence of missing them.
  • Failover design across sites or regions, sized for real load rather than for the diagram.
  • Runbooks, roles and a communication plan covering staff, customers and regulators during the event.
  • Exercise programme, findings and the evidence pack showing what was tested and what it proved.

What you receive

DeliverableWhat it contains
Impact analysisServices ranked by consequence of loss, each with the dependency chain that has to come back before it can.
Recovery targetsRPO and RTO per service, costed and signed off by the business owner, so any gap is accepted knowingly.
Failover runbooksStep-by-step recovery per service, naming roles rather than individuals, with the decision points marked.
Exercise reportWhat was tested, what took longer than target, and the remediation items with owners and dates against them.

Reference architecture

A reference, not a template. Your estate decides which parts apply and in what order they arrive.

Business continuity reference architecture: resilience, control and visibility layersResilience: Impact Analysis, RPO / RTO Targets, Dependencies. Control: Failover Design, Runbooks, Communication Plan. Visibility: Exercise Reporting, Readiness Dashboard, Audit EvidenceResilienceImpact AnalysisRPO / RTO TargetsDependenciesControlFailover DesignRunbooksCommunication PlanVisibilityExercise ReportingReadiness DashboardAudit Evidence
Business continuity reference architecture: resilience, control and visibility layers

How success is measured

Targets are agreed with you before the work starts, and reported against for its duration.

  • Measured recovery time in the most recent exercise against the agreed RTO for that service.
  • Data loss observed at failover against the agreed RPO, measured during the exercise rather than assumed.
  • Share of in-scope services exercised within the agreed cycle, and the age of the oldest untested runbook.

Questions we are asked

  • How often should we test?

    A realistic pattern is a desktop walkthrough a few times a year and a technical failover at least annually, with any service that changed materially tested after the change. Frequency matters less than honesty. An exercise designed to succeed teaches you nothing; the useful ones surface a broken step.

  • What RPO and RTO should we aim for?

    Those are business decisions with a price attached, not technical defaults. Near-zero data loss and minutes of downtime are achievable and cost accordingly. Four hours costs a fraction of that. The workable method is to state what each level costs, let the service owner choose, and then design to the choice.

  • Isn't cloud backup enough?

    Backup is one part of recovery and the part most often mistaken for all of it. A backup tells you the data still exists. It does not tell you how long a restore takes, whether the target environment exists, or in what order services must return. Untested restores are the most common gap we find.

  • Does replication protect us from ransomware?

    Not by itself, and assuming otherwise is dangerous, because replication faithfully copies encrypted data to the secondary site. That case needs immutable or otherwise isolated copies, with a retention window longer than the time it typically takes to notice an intrusion. It is a different design from availability replication, and both are usually needed.

  • Can our secondary site actually run production?

    That is precisely the question an exercise answers, and the answer is often at reduced capacity. A secondary sized for a subset of load is a legitimate choice when it is a stated one, because then the plan says which services get shed. The problem is a secondary assumed to be equal that has never carried full load.

  • Who should declare a disaster?

    A named role, with the authority written down in advance and a deputy for when that person is unreachable. The expensive delays in real incidents are usually decision delays rather than technical ones. The plan should also state what evidence triggers the call, so the threshold is not being invented while the clock runs.

Continue reading

  • Cloud

    The full domain, and the other capabilities within it.

  • Cloud Migration

    Application inventory, landing zone design and wave-based cutover for moving workloads to AWS, Azure or GCP with dependencies mapped before the window.

  • Hybrid Cloud

    Landing zone, private connectivity and one identity and policy model spanning on-premises and cloud, so workload placement becomes a decision.

Start with an assessment

The fastest way to a useful answer is a short, scoped look at what you already have.