drj logo
drj logo

Welcome to DRJ

Already registered user? Please login here

Create new account
(it's completely free). Subscribe

x

When the Playbook Fails: Why Traditional BCP Frameworks Weren’t Built for AI

AI: Automation & InnovationBusiness Continuity ManagementCyber Resilience & IT Disaster RecoveryEmerging Threats: Geopolitical & Climate RiskFeaturedGovernance: Compliance & Regulatory ReadinessOperational ResilienceResilience Strategy & Program MaturityRisk Management & Quantification
When the Playbook Fails: Why Traditional BCP Frameworks Weren't Built for AI

It starts quietly.

An AI system deployed across lending operations begins flagging applications using a pattern which made sense in testing but drifts from reality in production. No alarm fires. No error message surfaces. The system is performing exactly as designed.

By the time a human reviews the output, thousands of decisions have been made, and the downstream impact is already moving through the organization at a speed no incident response team was built to match.

But here's the part that doesn't make headlines: when the organization tries to revert to manual processing, they discover the team that used to run that process has been restructured. The institutional knowledge that lived in those people walked out the door. The backup isn't a process. It was a person. And that person is gone.

This isn't hypothetical. It's the shape of risk business continuity professionals are increasingly being asked to manage, using frameworks that were never designed for it.

What Traditional BCP Was Built to Handle

Business continuity planning as a discipline was built around a set of assumptions that were held for decades because they matched reality.

Failures were human scale. One system went down. One process broke. One bad decision was made by one person in one place. The impact was contained. The cause was traceable. The recovery path, while difficult, was linear.

Critically, recovery assumed humans were available to step in. The manual workaround existed. The staff who understood the process existed. The institutional knowledge needed to bridge the gap between disruption and restoration existed because humans had been doing the work before the technology failed.

Our frameworks reflect those assumptions. We identify critical processes. We map dependencies. We document recovery time objectives. We run tabletop exercises that walk teams through scenarios step by step, with clearly defined decision points and clear chains of accountability.

These are good practices. They remain essential. But they were designed for a world where failure moved at human speed and where humans were always available as the fallback. That world is changing faster than most continuity plans acknowledge.

How AI Changes Every Assumption

Artificial intelligence doesn't fail the way humans do. That statement sounds simple. Its implications are not.

When a human underwriter makes a flawed judgment call, the impact is singular. One loan. One relationship. One recoverable error. When an AI underwriting model applies that same flawed logic, it applies it to every application it processes simultaneously, continuously, and often invisibly until someone notices something is wrong.

By that point, the error isn't contained. It's replicated across the system at machine speed.

Traditional BCP assumes failures are traceable. With AI, failure may not look like a failure at all. The system is running. Outputs are being generated. Metrics look normal. The problem isn't the system stopped working. It continued working in a way that was quietly wrong.

Traditional BCP assumes accountability is clear. With AI, that clarity dissolves. When the decision-maker is an algorithm, the question of who owns the failure the team that deployed it, the vendor that built it, the leader who approved it becomes genuinely complicated, and organizations that haven't worked through this question in advance will find themselves working through it in the middle of a crisis.

Traditional BCP assumes recovery means restoration. With AI, restoration raises its own questions. What does “returning to normal operations” mean when normal included an autonomous process making decisions your team may not fully understand?

Perhaps most critically, traditional BCP assumes when technology fails, people can take over.

That assumption is now at risk.

The Four Gaps Most Organizations Haven't Closed

Gap 1: Detection

In traditional continuity planning, detection is relatively straightforward. A system goes offline. A process fails. An alert fires. Someone responds.

AI failures are often silent. The system continues operating. Outputs continue generating. The drift from expected behavior can be gradual enough no single data point triggers concern until the cumulative impact becomes impossible to ignore.

Closing this gap requires rethinking what monitoring means. It's not enough to know the system is running. Organizations need visibility into whether it's running correctly and that requires defining what "correctly" looks like before deployment, not after something goes wrong.

Gap 2: Accountability

Every traditional BCP has a clear answer to the question: who is responsible when something breaks? There is a chain of command. There are defined roles. There are decision rights.

AI complicates this in ways organizational charts weren't designed to handle. When an agentic AI process fails mid-execution or when it completes successfully but produces harmful outcomes accountability can become genuinely ambiguous.

Organizations that close this gap do so by assigning human accountability for AI system outcomes before deployment. Not after. The humans responsible for the consequences of AI decisions need to be identified, empowered, and trained before the system goes live, not introduced to the problem when something has already gone wrong.

Gap 3: Recovery

Recovery in traditional BCP means restoring the process to its pre-disruption state. The backup kicks in. The system comes back online. Operations resume.

Recovery from an AI-related disruption is more complex. If the failure was the result of model drift, a gradual degradation in the quality of AI outputs, returning to the previous state means returning to the conditions which produced the failure. If the failure was the result of a flawed training dataset, recovery requires understanding and correcting the flaw before the system can be trusted again.

These are not technical problems alone. They are operational problems, governance problems, and leadership problems. They require continuity professionals to be in the room when AI systems are designed and deployed, not called in after the incident has already occurred.

Gap 4: The Vanishing Fallback

This is the gap nobody is talking about, and it may be the most consequential one.

For generations, the human fallback was the foundation of every business continuity strategy. When technology failed, people stepped in. They knew the process. They had done it before. Institutional knowledge existed in the organization because the work had been done by people before it was automated.

AI adoption is systematically dismantling that foundation.

As organizations deploy AI to replace or reduce human-performed functions, they are also, often intentionally, sometimes invisibly, eliminating the very resources on which a traditional recovery strategy depends. Roles are restructured. Headcount is reduced. Subject matter experts are redeployed or not replaced when they leave. The institutional knowledge that once lived in people is allowed to quietly expire because AI has made it seem unnecessary.

Then the AI goes down.

The organization discovers it no longer has the staff to run the process manually. The people who understood the nuances, the edge cases, the judgment calls, the exceptions that never made it into the training data are gone. What remains is a recovery plan that points to a human fallback that no longer exists.

This is not a theoretical future risk. It is happening in organizations right now, in real time, as AI adoption accelerates and workforce decisions follow close behind. The gap between "AI handles this" and "someone on our team understands this well enough to do it manually" is widening with every restructuring, every voluntary departure that goes unbackfilled, every training program that stops teaching skills the AI has taken over.

The longer this continues, the harder it becomes to reverse. Institutional knowledge doesn't just pause while AI runs the process. It atrophies. And once a critical mass of subject matter expertise has left the organization, rebuilding it isn't a matter of weeks. It's a matter of years if it's possible at all.

What Needs to Change

The good news is that closing these gaps doesn't require rebuilding business continuity from the ground up. It requires extending frameworks that already work to cover a category of risk they weren't originally designed for.

Here are five places to start.

Stress-test your AI assumptions

Most organizations have not formally asked: what are we assuming about our AI systems that could be wrong? Start there. Map the assumptions embedded in each AI-enabled process and pressure-test them the way you would any other critical dependency.

Build AI-specific failure scenarios

Tabletop exercises that adapt existing scenarios to include an AI component are a good starting point but they're not sufficient. Organizations need failure scenarios that are native to AI risk: silent drift, cascading autonomous decisions, accountability gaps, and recovery paths that don't map cleanly to existing playbooks.

Embed human oversight before deployment

The instinct is to add oversight after something goes wrong. Resist it. Organizations managing AI risk most effectively are the ones that define human checkpoint requirements as a condition of deployment before the system goes live, not after the first incident.

Protect institutional knowledge as a strategic asset

Before any AI system is deployed to replace or augment a human-performed function, organizations should formally document the knowledge, judgment, and expertise held by the people currently doing the work. Not just the process steps but the edge cases, the exceptions, the nuances experienced practitioners carry in ways that are rarely written down. That knowledge is a continuity asset. Treat it like one.

Audit your manual fallback capability honestly

For every AI-enabled critical process, ask the hard question: if this system went offline today, could our current team run this process manually? Not in theory. In practice. With the staff and expertise that actually exist in the organization right now. If the answer is no or not for long, that is a continuity risk that belongs in your risk register, regardless of how reliable the AI system appears to be.

The Playbook Needs a New Chapter

Business continuity planning has always been an act of imagination, the discipline of preparing for what hasn't happened yet. Professionals who do this work well are the ones who resist the assumption yesterday's risks define tomorrow's threats.

AI represents a genuine expansion of the risk landscape. Not a replacement of what came before, but an addition. One that moves faster, fails differently, and is quietly reshaping the human infrastructure on which recovery has always depended.

Organizations that will navigate this well are the ones that treat AI resilience not as a technology problem, but as an organizational one. That means keeping humans close to the processes AI is running. It means protecting the expertise that makes manual recovery possible. It means asking hard questions about workforce decisions before the AI goes down not after.

The playbook has always been a living document.

It's time to add a new chapter. And this time, the most important section isn't about the technology.

It's about the people.

ABOUT THE AUTHOR

Abbey Hernandez

Abbey Hernandez is business resiliency director at a global financial services institution, where she leads the strategy and programs responsible for keeping one of the world's largest financial institutions operational when the unexpected happens. With more than 20 years in financial services, an MBA from the University of Wisconsin-Milwaukee, and certificates in executive women in leadership and ai strategy from Cornell University, Hernandez writes and speaks about what organizations must do differently to build resilience in an AI-driven world. The opinions expressed in this article are the author’s own and do not represent the views of her employer.

Latest News
DRJ HOT ITEMS
Webinar Spotlight
Fetching Upcoming Webinars...
Journal Categories

AI: Automation & Innovation

Business Continuity Management

Crisis Management & Emergency Response

Cyber Resilience & IT Disaster Recovery

Leadership: Culture & Workforce Resilience

Operational Resilience

Risk Management & Quantification

Sector-Specific & Critical Infrastructure Resilience

Supply Chain & Third-Party Resilience

Governance: Compliance & Regulatory Readiness

Incident Management & Response Coordination

Resilience Strategy & Program Maturity

Data Protection: Backup & Recovery

Exercises: Testing & Scenario Planning

Emerging Threats: Geopolitical & Climate Risk

Contact Us

Newsletter

The Journal, right in your inbox.