Risk

Business Continuity and Operational Resilience: A Career Path

Operational resilience turned business continuity from a once-a-year plan-writing exercise into a regulated discipline with its own career ladder. Here is what the work involves, who hires for it, and how to get in.

Two-color print illustration of a balance scale weighing a stack of documents against a certificate with a wax seal.

A regional bank's core payments processor goes down for four hours on a Tuesday morning. Nobody dies, no data is stolen, and by lunchtime the systems are back. But under the UK's operational resilience rules, and increasingly under equivalent regimes elsewhere, that outage is not just an IT incident. It is a test of whether the bank had already mapped this exact failure mode against its "important business services," pre-defined an impact tolerance for how long payments could be down before customers were harmed, and rehearsed who does what in the first thirty minutes. The team that built that map, set that tolerance, and ran that rehearsal is the operational resilience function — and it did not exist in most firms a decade ago.

Business continuity used to mean a binder on a shelf: a plan, refreshed once a year, that named a backup office and a call tree. Operational resilience absorbed that discipline and gave it teeth, a budget, and a reporting line to the board. For people who like risk work with a clear before-and-after test — did the organization actually survive the disruption, not just document that it planned to — it has become one of the more concrete corners of the risk profession to build a career in.

What changed, and why the jobs multiplied

Three regulatory moves did most of the work. The UK's PRA and FCA rules (effective 2022, with firms required to be within impact tolerances by March 2025) require banks, insurers, and financial market infrastructure firms to identify their important business services, set impact tolerances for each, and prove through testing that they can stay within them. The EU's Digital Operational Resilience Act (DORA), which applies from January 2025, does something similar across the EU financial sector with a heavier focus on ICT and third-party technology risk. In the US, the FFIEC's interagency guidance on operational resilience and existing business continuity examination standards push banks toward the same outcome without a single unified statute.

The effect across all three regimes is similar: firms had to stop treating continuity as an IT recovery exercise and start treating it as an enterprise discipline that spans technology, people, premises, third parties, and the business services those things support together. That reframing created new roles, because a plan that only IT understands cannot satisfy a regulator asking whether the firm can actually keep processing customer payments during a data center failure.

What the work actually involves

The center of the discipline is the business impact analysis (BIA): systematically working out which services matter most, what depends on what, and how long each dependency can be broken before the damage becomes unacceptable. From there, the job branches into several recurring workstreams.

  • **Mapping important business services end to end.** For each critical service, an analyst traces every process, system, third party, and person it depends on. A "process a mortgage payment" service might depend on a core banking platform, a specific data center, a payments network, three internal teams, and two outsourced vendors — and the resilience of the service is only as good as the weakest link in that chain.
  • **Setting and defending impact tolerances.** How many hours can a service be unavailable before it causes intolerable harm to customers or market stability? This is not a technical estimate; it is a judgment call that has to be justified to a board and, eventually, a regulator, using data on customer impact, financial loss, and reputational exposure.
  • **Scenario testing and severe-but-plausible exercises.** Teams design and run exercises — a cyberattack on a key vendor, loss of a primary site, failure of a critical third-party supplier — and test whether the organization actually stays within its stated tolerances. Tabletop exercises walk leadership through a scenario verbally; live exercises actually fail over a system or invoke a recovery site.
  • **Third-party and concentration risk.** A growing share of the work is understanding what happens when a critical vendor — a cloud provider, a payment processor, a data center operator — fails, and whether the firm has a workable exit or substitution plan rather than a contractual promise that assumes it away.
  • **Incident and crisis management.** When disruption actually happens, resilience teams often run or support the command structure that coordinates the response, then run the post-incident review that feeds lessons back into the BIA and the testing plan.
  • **Regulatory reporting and self-assessment.** Firms under formal regimes like the UK's must produce a self-assessment showing they understand their resilience posture and a board-approved statement of whether they can stay within tolerances. Someone has to assemble that evidence and defend it under challenge.

The skills that separate a strong analyst from a plan-writer

The field still has a reputation problem: some organizations run it as compliance paperwork, and some people in the role never get past writing plans nobody tests. The analysts who build real careers here share a different set of habits.

  • Comfort with systems thinking — the ability to trace a dependency chain across technology, vendors, and people rather than stopping at "the app is down."
  • Willingness to ask an uncomfortable question in a tabletop exercise: what if the backup site also depends on the thing that just failed? Weak exercises get designed to succeed; strong ones get designed to find the gap.
  • Data literacy to translate an impact tolerance into something measurable, not just a paragraph of prose that no test can confirm or refute.
  • The same escalation instinct that shows up across risk and audit work: recognizing when a gap is severe enough to go to senior management now, rather than waiting for the next scheduled report.
  • Enough technical fluency in infrastructure and third-party contracts to challenge a vendor's recovery claims instead of accepting them on faith.

How people actually get into it

There is no single feeder path, which makes this a relatively accessible specialty to move into laterally.

Business continuity and crisis management teams hire directly from operations, IT service management, and physical security backgrounds — people who already understand how the organization actually runs day to day. Internal audit and risk professionals move in because BIA and scenario testing use the same evidence discipline as a control walkthrough, just applied to disruption scenarios instead of process controls. IT and infrastructure staff move in from disaster recovery and incident response roles, bringing the technical depth that mapping a dependency chain requires. And in regulated financial firms, some people move in from regulatory affairs or governance roles, drawn by the fact that this is one of the few risk disciplines with a hard, dated compliance deadline attached to it.

The realistic entry point for someone outside the field is usually a business continuity analyst or resilience analyst role focused on one part of the puzzle — running BIAs for a single division, coordinating a testing calendar, or maintaining the third-party inventory — rather than owning enterprise-wide impact tolerances from day one.

Certifications worth having

Two credentialing bodies dominate: the Business Continuity Institute (BCI), which offers the Certificate of the BCI (CBCI) as an entry credential, and DRI International, whose Associate Business Continuity Professional (ABCP) and Certified Business Continuity Professional (CBCP) track experienced practitioners toward more senior recognition. Both organizations also offer specialist and executive-level credentials for people running enterprise programs. ISO 22301, the international standard for business continuity management systems, is worth knowing even without formal auditor certification, because many programs are built directly against its structure and firms increasingly seek external certification against it.

None of these credentials substitute for having actually run a BIA or an exercise that found a real gap. Hiring managers in this space tend to ask for a specific example of a test that failed and what changed afterward — a portfolio answer, not a certificate list.

Where the career goes

An analyst who is good at BIAs and testing typically moves toward owning a business line's or region's resilience program within two to three years, taking responsibility for its impact tolerances and its regulatory self-assessment. From there, the path splits much like other risk specialties: some move into enterprise resilience leadership, running the framework, the testing calendar, and the board reporting; others specialize further into third-party and technology resilience, a track that is getting more valuable as regimes like DORA push concentration risk with cloud and critical vendors to the top of supervisory attention. A smaller group moves sideways into crisis and incident management leadership, where the same dependency-mapping instinct applies in real time instead of in a planned exercise.

The discipline's appeal is also its discomfort: unlike a lot of risk work, operational resilience gets tested by reality on a schedule you do not control. A plan that only looks good on paper eventually meets an actual outage, and the gap between the two is the clearest measure of whether the work was real.