Business continuity plan template: what makes one work
Almost every organisation that has a business continuity plan has one that has never been used. That is not a criticism — the whole point is that you rarely need it. But it means most plans have never been tested against the only thing that matters, which is whether someone who did not write it can execute it while under pressure and short of options.
The plans that hold up share a small number of characteristics, and they are not the ones templates usually emphasise. They are readable by a stranger. They contain decisions made in advance rather than principles to be applied later. They are stored somewhere the disruption cannot reach. And their recovery objectives are numbers someone has actually measured, not numbers someone would like to be true.
This article walks through what goes into a plan that works, in the order you would build it. It applies whether you are writing one because a customer asked, because you are pursuing certification, or because NIS-2 now requires it — and there is a section on what that last case adds.
Short on time? The NIS-2 Compliance Suite contains this plan written out, along with the impact analysis, the crisis contacts sheet and the exercise report. €590 excl. VAT. If you first want to see where you stand, the free gap assessment takes ten minutes.
What the plan has to do
Before the structure, four jobs. If your plan does not do these, the rest is decoration.
It has to be reachable when everything else is not. A plan stored only on the file server that ransomware just encrypted does not exist. Keep an offline copy, a printed copy in a known place, or a copy in a second cloud tenant with separate credentials — and make sure someone other than its author knows where.
It has to be executable by whoever is available. The person who wrote it will be on holiday, or unreachable, or the incident will have started at two in the morning. Write for a competent colleague who has never read it.
It has to contain decisions, not principles. “Assess the impact and prioritise accordingly” is not a decision. “Restore the order system first, then the warehouse interface, then reporting” is. Every judgement you postpone to the moment of crisis is a judgement someone will make badly and at speed.
And it has to be honest about what you can actually do. A plan claiming four-hour recovery when nobody has ever restored anything faster than two days is worse than no plan, because it will be relied on.
Start with the impact analysis
RPO looks backwards at data loss. RTO and MTPD look forwards at downtime. All three are per service.
Everything in the plan derives from one question asked service by service: how long can this be unavailable before the consequences become unacceptable?
That produces three numbers, and the relationship between them is the backbone of the whole document.
MTPD — the maximum tolerable period of disruption — is the point beyond which the damage stops being recoverable. Not inconvenient: unacceptable. Contracts breached, customers gone, regulatory consequences triggered.
RTO — the recovery time objective — is how quickly you commit to restoring the service. It has to sit inside the MTPD with margin, because recovery never goes exactly as planned.
RPO — the recovery point objective — is how much data you accept losing, expressed in time. An RPO of four hours means you accept losing up to four hours of work, and it dictates how often you back up. This is the number people set carelessly and regret specifically.
Work through each critical service and record how the impact grows: after one hour, after a day, after a week. That progression is what tells you whether a service genuinely needs a two-hour RTO or whether someone simply said two hours because it sounded responsible.
| Service A | Service B | |
|---|---|---|
| Impact after 1 hour | Orders queue, recoverable | Nothing visible |
| Impact after 1 day | Contractual penalties begin | Reporting delayed |
| Impact after 1 week | Customers move to competitors | Month-end at risk |
| MTPD | 48 hours | 2 weeks |
| RTO | 8 hours | 3 days |
| RPO | 1 hour | 24 hours |
| Measured at last test | 11 hours | not tested |
That last row is the one most impact analyses do not have, and it is the one that makes the document true. The gap between the objective and the measured capability is either your improvement plan or your risk acceptance — but it has to be visible either way.
While you are there, record the dependencies for each service: the systems it needs, the people who know how it works, the sites and utilities, and the suppliers. Recovery order is usually dictated by dependencies rather than by importance, and finding that out during an incident is expensive.
The recovery strategy
With the numbers agreed, the plan can say what actually happens.
Start with activation. Who can invoke the plan, on what criteria, and how the team is convened. Name people, with deputies — a role title alone does not answer a phone. And specify the channel: if the team convenes on the corporate messaging platform, and the incident is a compromise of the identity provider that platform authenticates against, the plan has just failed at step one. An independent channel, tested, is worth more than several pages of technical procedure.
Then the sequence. Not a list of tasks but an ordered one, with the dependencies that force the order stated alongside. Database before application, network before both, authentication before anything that uses it. Where two things can happen in parallel, say so, because parallelism is where you recover the time the plan assumes.
Then the degraded mode, which most plans omit entirely. Between the incident and full recovery there is a period — often the longest part — where the service runs reduced or not at all. What happens in that window? Can orders be taken on paper? Can a subset of customers be served? What are the others told, and by whom? An organisation that has thought about this looks competent during a disruption; one that has not looks paralysed, and the difference is visible to customers.
Finally, the scenarios. A plan built around a single “disaster” scenario is fragile, because real disruptions are specific. Cover at least the loss of a critical system, the loss of data integrity including ransomware, the loss of a site, the loss of a critical supplier or cloud service, the loss of key people, and the loss of a supporting utility. Each of these changes what you do, and only one of them is the scenario most templates imagine.
Backup, and the part everyone gets wrong
A completed backup job proves data was written. It proves nothing about reading it back.
Your backup policy needs the ordinary things: scope, frequency aligned to the RPO, retention, encryption, and geographic separation from production. Then it needs the two that decide whether any of it works.
The first is an isolated copy. At least one backup has to be beyond the reach of the production environment and of the credentials that administer it — offline, immutable, or in a separate account with separate authentication. Ransomware that reaches your backups turns a bad week into an existential one, and it reaches them through the same administrative credentials that manage everything else. Replication provided by a cloud platform is not a backup: it faithfully replicates the encryption too.
The second is testing the restore rather than the backup. A backup job reporting success proves data was written. It says nothing about whether it can be read back, whether the restore tooling still works, whether anyone remembers the procedure, or how long it takes.
Test restores on a schedule, and record three things: what you restored, how long it took, and whether it worked completely. That third one matters — a restore that recovers the database but not the configuration is a partial restore, and finding that out during an incident costs you the whole RTO.
If your measured restore time is longer than your RTO, you have not failed. You have discovered something, which is the point of testing. Either the objective moves or the capability does, but the number in the plan should be one you have seen with your own eyes.
Writing this from scratch takes longer than most people budget for. The continuity procedure in the NIS-2 Compliance Suite covers impact analysis, recovery objectives, backup, disaster recovery, crisis management and exercising in one document — with the plan template, the crisis contacts sheet and the exercise report alongside it.
Crisis management and communication
Business continuity is about services. Crisis management is about the decisions and the communication that surround a disruption large enough that normal management does not cover it — and the two are separate capabilities that belong in the same document.
Define when a disruption becomes a crisis, who is on the crisis team, and who leads it. Then define who speaks to whom.
Communication is where organisations do the most damage to themselves. A few rules make it survivable.
One spokesperson, and nobody else speaks externally — including on social media, where an employee’s reassuring post can contradict the official position within minutes. Say what is known, what is not yet known, and when the next update will come; the third of those buys you more patience than the first two. Never speculate on cause or attribution before the investigation supports it, because retracting an attribution is worse than not having made one.
And keep what you tell customers consistent with what you tell any regulator involved. Two different accounts of the same event circulating at once turns a technical problem into a credibility problem.
Prepare holding statements in advance for the obvious situations: disruption with unknown cause, confirmed incident, incident affecting customer data, recovery in progress. Writing them under pressure produces the wording you later regret.
The contacts sheet — crisis team, key suppliers, insurer, incident response retainer, relevant authorities — needs to exist offline, be verified quarterly, and be reachable by someone standing in a car park with a phone. That is the actual use case.
What NIS-2 adds
If your organisation is in scope for NIS-2, continuity stops being good practice and becomes a legal obligation with specific content.
Article 21(2)(c) of Directive (EU) 2022/2555 requires business continuity, including backup management and disaster recovery, and crisis management. That is the whole of the directive text on the subject — deliberately outcome-based, naming the objective rather than the method.
The detail sits elsewhere. For entities in the digital infrastructure and digital provider sectors, Commission Implementing Regulation (EU) 2024/2690 sets out continuity requirements in section 4 of its Annex, covering the continuity and disaster recovery plan, backup management, and crisis management. That regulation applies directly and identically across the Union with no national transposition, so where it applies it is the more specific document.
Three things change in practice when NIS-2 applies.
The plan has to be evidenced, not just written. What gets inspected is the impact analysis with real numbers, the restore test reports with dates and durations, the exercise records, and the approval by the management body. A plan with no evidence of having been tested is treated as a plan that has not been tested.
Crisis communication acquires a regulatory recipient. Alongside customers, a significant incident triggers Article 23 notification duties — an early warning within 24 hours, a notification within 72, a final report within a month — and the clock runs from awareness, not from resolution. Recovery does not pause it. Assign the notification to someone who is not doing technical recovery, or it will be late.
And exercises become non-optional. The measure that requires you to assess the effectiveness of your controls, Article 21(2)(f), applies to this one too. An untested plan is not an effective control, whatever it says on the cover.
Exercising the plan
An exercise everyone passes has measured nothing. The purpose is to find what does not work while it is cheap to find out.
A tabletop once a year is the minimum: pick an uncomfortable scenario, put the team in a room without access to their systems, and watch what happens. The failures tend to be mundane and revealing — the contact list only exists on the intranet, the person who knows the restore procedure left in March, nobody can authorise the spend on emergency hardware.
Technical restore tests should be more frequent than that, and at least one exercise a year should be unannounced if your organisation can bear it. Announced exercises measure the plan; unannounced ones measure the organisation.
Record what happened, including the parts that went badly, and track the actions to closure. An exercise report with no actions is a report that has been written for the file rather than for the team.
The plan, section by section
Adapt the wording; keep the structure.
- Scope and objectives — which services this plan covers, and the MTPD, RTO and RPO for each.
- Dependencies — systems, people, sites, utilities, suppliers, and the workaround for each.
- Activation — who can invoke, on what criteria, and how the team is convened.
- Roles — names and deputies, not job titles alone.
- Communication channel — the one that works when the primary environment does not.
- Recovery sequence — ordered, with the dependencies that force the order and expected durations.
- Degraded mode — how the service continues, reduced, while recovery proceeds.
- Scenario notes — what changes for system loss, ransomware, site loss, supplier loss, people loss, utility loss.
- Access during recovery — credentials, where they are held, under what control.
- Crisis management — when a disruption becomes a crisis, who decides, who communicates.
- Holding statements — drafted in advance for the predictable situations.
- Completion criteria — how you decide recovery is over, and who declares it.
- Testing — schedule, scope, and where the reports live.
- Version control — date, owner, approval, and the date of the last exercise.
Four things to do this week
If writing the whole plan is more than you can take on now, these four give the most back for the least effort.
Find out when you last restored something, and how long it took. If the answer is “we have backups” rather than a date and a duration, that is the finding, and booking a test is the fix.
Check whether one backup copy is genuinely out of reach of your production credentials. If an administrator account can delete it, so can whoever compromises that account.
Print the contact list and put a copy somewhere that does not require a login. It costs five minutes and it is the single most likely thing to be missing when you need it.
Ask two people who are not you what they would do if the primary system were unavailable tomorrow morning. Their answers tell you more about the state of your continuity planning than the document does.
Frequently asked questions
What is the difference between a business continuity plan and a disaster recovery plan? Continuity is about keeping the business operating — including without IT, in degraded mode, by other means. Disaster recovery is the subset concerned with restoring the technology. Many organisations keep them as one document, which is fine as long as the non-technical half actually exists.
How do we set RTO and RPO? From the impact analysis, not from a wish. Work out how the consequences grow over time for each service, agree the point at which they become unacceptable, and set the objective inside that with margin. Then test, and record the number you actually achieved.
How often should the plan be tested? A tabletop at least annually and technical restore tests more frequently — quarterly is a common cadence for critical systems. Also test after any significant change to your services, systems or team.
Does NIS-2 require a business continuity plan? Article 21(2)(c) requires business continuity including backup management and disaster recovery, and crisis management. It does not prescribe a format, but it does require you to be able to evidence that what you have works, which in practice means impact analysis, tested restores and exercise records.
Who should own the plan? Someone senior enough to make the decisions it contains, with a named deputy. Ownership by “IT” as a department is how plans end up unmaintained, because a department cannot be asked when it was last reviewed.
Where to go next
Climate-related events belong in the threat analysis behind any continuity plan, and under ISO 27001 considering them is no longer optional: amendment A1:2024 requires the determination to be made and recorded, whichever way it goes.
If you are in scope for NIS-2, continuity is one of ten risk-management measures, and knowing where you stand on the other nine matters as much.
The free NIS-2 gap assessment asks twelve questions, one per obligation, and scores you measure by measure. Ten minutes, and it tells you which obligations you could evidence today.
If you would rather start from a written set, the NIS-2 Compliance Suite contains the continuity procedure described here — impact analysis, recovery objectives, backup, disaster recovery, crisis management and exercising in one document rather than seventeen — plus the crisis contacts sheet and the exercise report. €590 excl. VAT.
You may also want to read our guide to an incident response plan that survives the 24-hour clock, which covers the reporting duties that run alongside recovery, and the ten Article 21 measures.
Sources
Directive (EU) 2022/2555 and Commission Implementing Regulation (EU) 2024/2690 are available on EUR-Lex.