Business Continuity Plan: How to Write One That Works

Business Continuity Plan: How to Write One That Works

2026-09-10 · Tommaso Maria Ricci

Before you write a business continuity plan, look at what an interruption actually costs. More than half of the organizations surveyed by the Uptime Institute said their most recent major outage cost over one hundred thousand dollars, and for the second year running, one in five put the figure above a million. That is the 2025 Annual Outage Analysis, and it measures data center operators, the people who are supposed to be good at this.

Now consider the company that has never measured anything. The IBM 2025 Cost of a Data Breach Report found that 86% of breached organizations reported operational disruption: delayed sales, interrupted service, halted production. Not data loss. Work that stopped.

Most companies respond to this with a document. Someone writes a business continuity plan, usually because a customer asked for one during a procurement review, and the file goes into a folder. It is forty pages long, it lists phone numbers that are two job changes out of date, and nobody has read past page three since the day it was approved.

That document is worse than nothing, because it converts an open risk into a closed checkbox. This guide is about writing the other kind: short, tested, and owned by named people who know they own it. It covers what goes in a business continuity plan, how to size recovery targets so they mean something, what to test and how often, and what the plan should never try to be.

What a business continuity plan actually is

Strip away the vocabulary and a business continuity plan answers four questions.

What must keep running. Not everything. A short list of activities the business cannot survive without for long, ranked by how fast the damage compounds.

How fast each one must come back. Expressed as a number of hours, and derived from analysis rather than from optimism.

What we do in the first hours. Who declares an incident, who talks to customers, who talks to staff, what the fallback way of working looks like when the normal one is unavailable.

How we prove it works. A test schedule, with results, and a mechanism for fixing what the test broke.

Everything else in a typical plan is either supporting detail or padding. The international standard for this, ISO 22301, formalizes the same logic through a business impact analysis, a risk assessment, recovery strategies, and documented procedures, and the freely available NIST contingency planning guide, SP 800-34 walks through the same sequence without a paywall. You do not need certification to use the structure, and most mid-sized companies should not chase certification until a customer contract requires it.

What it is not

It is not a disaster recovery plan. Disaster recovery is the technical subset: restoring systems, data, and infrastructure. Continuity is the business question of how the company keeps serving customers while the technical people work. Companies that conflate the two end up with an excellent runbook for restoring a database and no answer for what the sales team tells customers on day two.

It is not a crisis communications plan either, though it must contain one. And it is not an insurance document. Insurance pays for losses after the fact. Continuity is about the losses that insurance never covers, which are the customers who left during the outage and did not come back.

Why most continuity plans fail before they are used

Having read a fair number of these documents inside real companies, the failure modes are remarkably consistent.

The plan was written by one person alone. Usually someone in IT or compliance, working from a template, without interviewing the people who actually run the operations being described. The result is technically literate and operationally fictional.

It assumes the disruption is total. Real disruptions are partial and ambiguous. The building is fine but the phone system is down. The system is up but the data is wrong. Half the team is unreachable and nobody is sure whether this counts as an incident yet. Plans written for the total-loss scenario give no guidance in the ambiguous ones, which are the ones that actually happen.

Recovery targets were chosen to sound reassuring. Somebody wrote four hours because four hours sounded serious without sounding expensive. No one costed what four hours actually requires, so the number is aspirational, and everybody involved quietly knows it.

It has never been tested against resistance. A tabletop exercise where everyone agrees the plan works is not a test. A test is when somebody plays the customer who is threatening to leave, and the person holding the plan discovers there is no script for that.

Nobody owns it between incidents. The plan has an author, which is not the same as an owner. Authors finish. Owners maintain, and maintenance is the entire job.

The business impact analysis, done in a week

This is the part everyone skips, and skipping it is why the rest of the plan floats free of reality. Done properly it does not take months. In a company under three hundred people it takes about a week of calendar time and maybe twenty hours of actual work.

Step one: list the activities, not the systems

Write down what the business does, in business language. Take customer orders. Produce and ship. Invoice and collect. Pay staff. Answer support requests. Maintain regulated records. Keep between eight and fifteen items. Anyone who lists forty has listed tasks, not activities.

Resist the pull toward systems. The system list comes later and it comes from this list, never the other way around. Companies that start from the system inventory end up protecting the systems that are easy to describe rather than the ones that matter.

Step two: ask how fast the damage compounds

For each activity, ask a single question of the person who runs it: if this stopped right now, when does it start to hurt in a way we cannot undo?

The answers cluster into useful bands. Some activities tolerate days without permanent harm. Some tolerate hours. A few tolerate almost nothing, usually because of a regulatory deadline, a perishable good, a safety implication, or a customer contract with a service credit attached.

Record the answer as a number, and record the reason next to it. The reason is what makes the number defensible later, when someone in finance asks why the recovery capability costs what it costs.

Step three: set recovery time and recovery point

Recovery time objective is how long an activity can stay down. Recovery point objective is how much data you can afford to lose, measured in time. They come out of the analysis above, and, as any ISO 22301 auditor will tell you, they must trace back to it rather than to a number that sounded good.

These two figures drive everything expensive: backup frequency, whether standby capacity is warranted, what your contracts with alternate suppliers need to say, how many people need cross-training. Set them casually and you will either overspend on the unimportant or discover the gap during the incident.

Step four: find the single points of failure

For each critical activity, list what it depends on: a system, a supplier, a physical location, a specific person. Then mark every dependency that has no alternative.

The dependency that surprises people is almost never technical. It is the one employee who knows how the pricing exception process works, the single supplier for a component with an eleven-week lead time, the bank relationship that only one director can authorize. The technical single points of failure are usually already known. The human and contractual ones are not, and they are the ones that turn a two-day problem into a two-month one. Companies that have built a proper vendor management program find this step much faster, because the supplier dependencies are already mapped.

How to write a business continuity plan that people will use

Here is the format that survives contact with an actual incident. It is short by design.

The first page is the only page that matters at hour zero

One page. Who declares an incident and by what authority. Their deputy, and the deputy's deputy. The three numbers to call. Where the team assembles, physically or virtually, and the fallback if that channel is the thing that is down. What gets communicated in the first sixty minutes, and to whom.

If your first page is a table of contents, rewrite it. Nobody reads a table of contents during an outage.

Then the activity playbooks

One page per critical activity, and no more than one. Each contains what the activity is, its recovery time objective, its dependencies, the manual or degraded way of doing it, who runs the degraded process, and what has to happen to declare it recovered.

The degraded process is the part most plans omit and the part that gets used most. If the order system is down, can orders be taken on paper and entered later, and does anyone know the format? If the payment system is down, is there an approved way to promise a customer something? These procedures need to exist on paper before they are needed, because inventing them under pressure is how companies make commitments they cannot keep.

Then the communications section

Pre-drafted, pre-approved messages for the four audiences: staff, customers, suppliers, and, where relevant, regulators. Not final text, which will always be wrong for the specific situation, but a skeleton with the decisions already made about tone, channel, and who signs.

Include the thing that trips up almost everyone: what you say when you do not know yet. The default corporate instinct is to say nothing until the picture is clear, and the default customer interpretation of silence is that you are hiding something worse.

Then the appendices, which nobody reads

Contact lists, system inventories, recovery runbooks, insurance details, supplier contracts. These belong at the back, and they belong in a form that can be updated without reissuing the plan. A contact list embedded in a PDF that requires a governance committee to revise is a contact list that will be wrong.

Sizing what it costs, honestly

Continuity capability costs money, and the argument for it fails when the numbers are vague on both sides.

The cost side has three components. Standby capacity, which is the expensive one: duplicate systems, alternate premises, retained suppliers, held inventory. Process work, which is the cheap one: writing degraded procedures, cross-training, documenting the undocumented. Testing, which is recurring and usually underestimated because it consumes the time of exactly the senior people who are hardest to free up.

The benefit side is harder and must be handled without exaggeration. The honest framing is not a probability-weighted expected loss, because nobody believes those numbers. The honest framing is a threshold question: what is the longest interruption this business could absorb and still be here afterward, and are we currently able to stay inside it?

The asymmetry that justifies the cheap half

Here is what makes this decision easier than it looks. Most of the value comes from the cheap half.

Writing down the degraded process for taking orders costs a day of somebody's time. Cross-training a second person on the pricing exception process costs a week. Documenting which supplier has an eleven-week lead time costs an afternoon. These items cover the majority of realistic scenarios and cost almost nothing, and they are consistently deferred in favor of debating whether to fund a secondary site.

Do the cheap half first, completely, before costing the expensive half at all. Companies that reverse this order spend a year in a capital debate and remain undefended against the disruption that was going to hit them anyway. The same sequencing logic applies when deciding what to automate, which is covered in the guide to business process automation.

Testing: the part that separates a plan from a document

An untested plan is a hypothesis. Testing is the only mechanism that converts it into a capability, and the testing schedule is the single best predictor of whether a company will cope.

There are four levels, and they escalate in cost and value.

Read and confirm. Each owner reads their own page and confirms it is still accurate. Costs an hour per quarter. Catches the decay that happens silently: people who left, systems that were replaced, suppliers that changed.

Tabletop. The team walks through a scenario in a room, talking. Two hours, twice a year. Its real value is not confirming the plan works. It is exposing the disagreements about who decides what, which surface immediately and never surface any other way.

Functional test. One component is actually exercised. Restore a backup and time it. Run the order process on the degraded path for half a day. Call the alternate supplier and ask them to confirm, in writing, what they could actually deliver on short notice. This is where comfortable assumptions die.

Full simulation. A scenario is run end to end with the systems genuinely unavailable. Expensive, disruptive, and appropriate roughly once every two years for most mid-sized companies, more often in regulated sectors.

The test that everyone skips and everyone needs

Restore something from backup, at random, and time it with a stopwatch.

Not verify that the backup job completed. Restore an actual file, or better, an actual system, and measure how long it took from the decision to the moment someone could use it. In organizations that have never done this, the measured time exceeds the assumed time by a wide margin, and occasionally the restore fails outright because nobody had ever exercised it.

This single test, done once a year and recorded, is worth more than the next twenty pages of plan.

Record what broke

Every test produces findings. The findings need an owner, a date, and a place where somebody looks at them. A test whose findings go nowhere has trained the organization that testing is theater, which is worse than not testing at all, because it inoculates people against taking the next one seriously.

Where regulation forces your hand

For a growing set of companies, continuity is no longer optional and the requirement comes with an auditor attached.

Financial services in the EU. The Digital Operational Resilience Act has applied since 17 January 2025, and it requires financial entities to maintain ICT risk management, incident reporting, and a testing program, including threat-led penetration testing for the entities in scope. It also reaches third-party ICT providers, which means firms outside financial services can find themselves inside the perimeter through a contract.

Essential and important entities under NIS2. The EU network and information security directive requires business continuity, backup management, and crisis management among its risk management measures, and it extends to supply chain security. Member state implementation timelines vary, so the practical question is not whether the directive applies to your sector in the abstract but what your national transposition requires and by when.

Contractual requirements. Increasingly the binding constraint arrives through procurement rather than legislation. Large customers ask for evidence of a tested plan, and the evidence they want is the test record, not the document. Companies that treat continuity as a sales enabler rather than a compliance chore tend to build the better version, because they have to show it to someone skeptical.

The overlap between these regimes is substantial, which is good news: one well-built continuity capability answers most of them. Where they differ is mainly in reporting deadlines and testing frequency. The broader compliance architecture around this is covered in the guide to AI governance for business, which deals with the same problem of building one control set that satisfies several regimes.

The scenarios worth planning for

Plans that enumerate every conceivable disaster are unusable. Plans built around a small number of loss types are usable, because the response depends on what you lost, not on what caused it.

Loss of premises. Fire, flood, structural, access denial. Question: where does the work happen tomorrow, and what physically cannot be replicated elsewhere?

Loss of systems or data. Failure, cyber incident, supplier outage, corruption. Question: what is the manual path, and how long can it hold?

Loss of people. Illness, a resignation at a bad moment, an accident, a strike. Question: who else can do this, and have they ever actually done it?

Loss of a supplier. Insolvency, force majeure, a dispute that escalates. Question: what is the lead time to switch, and does the contract permit it? Answering the second half requires knowing what your agreements actually say, which is a contract lifecycle management problem before it is a continuity one.

Loss of reputation or trust. A public incident, a regulatory action, a viral complaint. Question: who speaks, and what do they say in the first hour?

Four or five scenarios covering these categories will exercise nearly every part of a plan, and the public preparedness material from Ready.gov is a reasonable sanity check on whether you have missed a category. Ransomware deserves separate attention only because it combines two categories at once, systems and trust, and because the decision structure it forces is unlike the others. The technical dimension is covered in the guide to AI for cybersecurity in business.

A case from the field: a medical practice that found its real constraint

A medical center I worked with came to the problem sideways. The stated goal was capacity: they wanted to see more patients without adding staff or hours, and the work delivered roughly twenty percent additional capacity through scheduling and process changes.

The continuity issue surfaced during the analysis, not because anyone had asked. Mapping the appointment and records workflow made something obvious that had never been written down: if the practice management system was unavailable for a full day, there was no way to know who was expected to arrive, no way to access clinical history, and no agreed way to decide which patients could safely be seen anyway and which had to be rescheduled.

Nobody had considered this a risk because it had never happened. The fix was not technical and did not require a purchase. It was a printed next-day schedule generated automatically each evening, a documented rule for which appointment types could proceed without record access and which could not, and one named person authorized to make the call. Total cost: about half a day of work, and a recurring print job.

Some months later a regional connectivity failure took the system offline for most of a morning. The practice ran on the printed list, rescheduled the appointments the rule said to reschedule, and lost a fraction of the day rather than all of it.

The transferable point is that the entire benefit came from the cheap half. There was no standby system, no secondary site, no vendor contract. There was a piece of paper, a rule, and a name.

If you suspect your own operation has a dependency like that one, the fastest way to find it is not a risk register. It is walking one critical process end to end and asking, at each step, what happens if this specific thing is unavailable right now.

Self-assessment scorecard

Answer yes or no. Every no is a work item, not a verdict.

Analysis

  1. Is there a written list of critical business activities, ranked, updated within the last twelve months?
  2. Does each critical activity have a recovery time objective with a documented reason behind the number?
  3. Have single points of failure been identified for people and suppliers, not only for systems?
  4. Do you know which single supplier failure would take longest to work around?

Plan

  1. Does the plan fit on a page for hour zero, with named individuals and deputies?
  2. Does each critical activity have a documented degraded or manual procedure?
  3. Are there pre-approved communication skeletons for staff, customers, and suppliers?
  4. Is it clear who has the authority to declare an incident, and who acts if they are unreachable?

Testing

  1. Has a real restore from backup been performed and timed in the last twelve months?
  2. Has a tabletop exercise been run in the last twelve months with findings recorded?
  3. Do test findings have owners and due dates that somebody reviews?
  4. Has any alternate supplier confirmed in writing what they could actually supply at short notice?

Maintenance

  1. Is there a named owner of the plan, distinct from its original author?
  2. Is the contact list updateable without a formal reissue of the document?
  3. Does the plan get reviewed when the business changes, not only on a calendar cycle?
  4. Do the people named in the plan know they are named in it?

Reading the result

  • Fourteen or more yes: mature. The useful work is refining recovery targets and increasing test realism.
  • Nine to thirteen: solid foundation, gaps concentrated in testing or maintenance. Six months of ordered work.
  • Five to eight: you have a document, not a capability. Start with the business impact analysis.
  • Four or fewer: the risk currently sits with whoever happens to be available on the day. Start this week, with question sixteen.

That last item is not filler. In more organizations than you would expect, people discover during an actual incident that they were listed as the deputy incident commander in a document they had never seen.

Roadmap for 30, 60 and 90 days

First 30 days: understand

The goal is not to write anything polished. It is to stop guessing.

  • Interview the owner of each business function and produce the list of critical activities, eight to fifteen items
  • For each, capture the tolerable outage duration and the reason behind it
  • Map dependencies per activity, marking every one that has no alternative
  • Identify the three dependencies with no alternative and the longest replacement time
  • Perform one real restore from backup and record the elapsed time

The last item routinely produces the most uncomfortable finding of the quarter, and it is far better to have it now, under controlled conditions.

Days 31 to 60: write and staff

  • Draft the one-page hour-zero page and get it approved by whoever has the authority to approve it
  • Write one page per critical activity, authored by the person who runs that activity rather than by a central function
  • Write the degraded procedures, which is where most of the real work sits
  • Draft communication skeletons for the four audiences
  • Cross-train a second person on the two or three tasks currently held by one individual

Days 61 to 90: test and hand over

  • Run a tabletop exercise on one realistic partial scenario, not a total-loss one
  • Record findings with owners and due dates
  • Contact the two most critical suppliers and obtain a written statement of what they could supply at short notice
  • Name the plan owner and give them scheduled time, not a title
  • Set the recurring calendar: quarterly read and confirm, half-yearly tabletop, annual restore test

Only after this cycle completes does it make sense to evaluate whether standby capacity is worth funding. Doing it earlier means costing a solution before you have measured the problem, and that debate consumes months.

Who owns this inside the company

Below about fifty employees the realistic answer is one person with a written mandate and one protected hour per week. Not a committee, not a project, and not a permanent external engagement.

Above fifty, three distinct roles are needed, and none of them requires a new function:

  • Someone who decides during an incident, with pre-agreed authority and a named deputy.
  • Someone who maintains the plan between incidents, with real allocated time rather than an addition to an existing full workload.
  • Someone who verifies, meaning schedules the tests and chases the findings, and who is not the same person as the maintainer.

Where continuity programs decay, the pattern is almost always identical: the maintainer exists but has no time, and the verifier does not exist at all.

There is also a cultural dimension that outweighs the documentation. In organizations where raising a concern is treated as pessimism, the risks that matter never reach the people who could fund a fix. The single points of failure are known, but they are known individually rather than institutionally. Getting that surfaced is a change management problem before it is a continuity problem.

What to expect, and when

The first benefit is not resilience. It is knowing what would actually happen, which changes decisions well before any incident occurs. Companies that complete the analysis usually make two or three unrelated operational changes in the first quarter simply because the mapping made something visible.

For a company of around one hundred people, the ninety-day cycle takes roughly sixty to ninety hours in total, spread across many people, with the heaviest load falling on the function owners writing their own pages. Ongoing maintenance settles at something like two hours a week plus the scheduled tests.

The measurable outcome that arrives first is usually the restore time, because it is the only figure that was previously assumed rather than measured. The second is a shorter list of single points of failure, achieved through cross-training rather than through spending.

If you are weighing whether to invest in standby capacity or to fix the process gaps first, that decision usually turns on three or four variables you already know, and a direct conversation about your specific dependencies tends to settle the sequence faster than another round of internal analysis.

FAQ

How to write a business continuity plan without producing another unread document?

Start from the business impact analysis rather than a template, and keep the plan short enough that people can hold it in their heads. The structure that survives real incidents is one page for hour zero with named individuals and their deputies, one page per critical activity containing its recovery target and its degraded manual procedure, a short communications section with pre-approved message skeletons, and appendices at the back that can be updated without reissuing the whole document. Have each activity page written by the person who actually runs that activity, not by a central function. A plan authored by one person working from a template is technically literate and operationally fictional, and it will be discovered as such at the worst moment.

What is the difference between a business continuity plan and a disaster recovery plan?

Disaster recovery is the technical subset: restoring systems, data, and infrastructure to working order. Business continuity is the wider question of how the organization keeps serving customers while that restoration is happening. Companies that conflate the two typically end up with an excellent database restore runbook and no answer for what the sales team tells customers on day two, or who is allowed to promise a delivery date when the planning system is unavailable. You need both, and the continuity plan should reference the recovery runbooks as appendices rather than containing them.

How do we set recovery time and recovery point objectives that are defensible?

Derive them from the impact analysis, and record the reason next to each number. Ask the person who runs each activity when a stoppage starts to cause damage that cannot be undone, and why: a regulatory deadline, a perishable good, a contractual service credit, a safety implication. That reason is what makes the figure survive scrutiny when finance asks why the capability costs what it costs. Targets chosen because they sound reassuring, four hours being the classic, drive spending decisions that nobody has actually costed, and any ISO 22301 auditor will expect the numbers to trace back to analysis rather than to instinct.

How often should a business continuity plan be tested?

Four cadences, escalating in cost. A quarterly read and confirm, where each owner checks their own page is still accurate, catches the silent decay of departed staff and replaced systems. A tabletop exercise twice a year exposes disagreements about decision authority that surface no other way. An annual functional test exercises one component for real, and the version that matters most is a timed restore from backup performed with a stopwatch. A full end-to-end simulation is appropriate roughly every two years for a mid-sized company, more often in regulated sectors. Every test must produce findings with owners and dates, otherwise it teaches the organization that testing is theater.

What does a continuity capability cost for a mid-sized company?

The cost splits into three parts, and they are wildly unequal. Standby capacity, meaning duplicate systems, alternate premises, or retained suppliers, is the expensive part. Process work, meaning writing degraded procedures, cross-training, and documenting the undocumented, is nearly free. Testing is recurring and is usually underestimated because it consumes senior people's time. The important asymmetry is that most of the protection comes from the cheap half. Complete the process work first, in full, before costing standby capacity at all. Companies that reverse the order spend a year in a capital debate while remaining undefended against the disruptions that were realistically going to happen.

Is a business continuity plan legally required?

It depends on sector and jurisdiction, and increasingly the binding requirement arrives through contracts rather than statute. Financial entities in the EU have been subject to the Digital Operational Resilience Act since 17 January 2025, which mandates ICT risk management, incident reporting, and a testing program, and which also reaches third-party ICT providers. Essential and important entities under the NIS2 directive face business continuity, backup, and crisis management obligations, with national transposition timelines that vary by member state. Outside regulated sectors, the practical trigger is usually a large customer asking for evidence during procurement, and the evidence they want is the test record rather than the document itself.

Do we need external help to build one?

For the analysis and the writing, usually not. Both require knowledge of how the business actually runs, which you have and an outsider does not. External help earns its cost in three specific situations: facilitating the tabletop exercise, because the person running the scenario needs to be willing to make participants uncomfortable and an internal colleague generally is not; validating recovery targets against what the technical capability can genuinely deliver, since those two figures are often set by different people who never compare notes; and preparing for a certification audit or a demanding customer assessment. For everything else the right person is already inside the company and simply lacks allocated time.