drj logo
drj logo

Welcome to DRJ

Already registered user? Please login here

Create new account
(it's completely free). Subscribe

x

Continuity of Operations Does Not Mean Business Continuity

168775In the field of emergency management, the concepts of continuity of operations (COOP) and business continuity management (BCM) are used interchangeably, as if they are two names for the same thing. They are not. There are some fundamental differences between the two operations which are worth examining for the sake of clarity, role definition, and planning methodology. It is not just semantics; there is a practical value in understanding whether an organization needs a COOP or BCM program. The difference can have some real impacts on the recoverability and resilience of the organization. 

COOP

It starts with the origins of the two disciplines. COOP originated with emergency management. Modern-day emergency management started as civil defense. Civil defense tried to protect Londoners from bombs in World War II, organizing shelters and enforcing blackouts. When the post-WWII Cold War period started, plans for protecting communities from disasters centered around another human-manufactured disaster: nuclear war. Fortunately for humanity, that threat has not been realized. Need proof? You are alive. 

The Cold War remained “cold” because the threat of a nuclear war kept the USSR bloc and the USA/Europe bloc from starting a new world war. For a decade from the mid-1950s, the US and the USSR depended on strategic bombers as the mechanism of their nuclear threat. Each side (and it is allies) built swarms of nuclear-capable long-range bombers that could reach their opponent’s territory and delivering their deadly payload. The advent of strategic intercontinental missiles changed the nature of the threatened conflict. In the era of the nuclear bomber, one could expect – rationally or otherwise –the US or the USSR could survive a nuclear exchange and remain a viable nation-state because only a limited number of bombers could penetrate the enemy’s air defenses. B-52 bombers (named after the year they came into service) and their Soviet counterparts could not be expected to deliver their payloads in amounts that would be world-ending. Guided nuclear missiles could not be shot down or stopped. They could deliver their payloads with a large degree of certainty. The introduction of the intercontinental ballistic missile (ICBM) was the end of civil defense.

There is no defense against ICBMs, and the use of those missiles would end human life on earth. This was referred to as “mutually assured destruction.”

When it became clear that nobody was going to survive a nuclear exchange, the concept of civil defense was soon dropped. Emergency managers started using the skills they learned in civil defense to address the disasters that occur regularly and with terrible impacts. Disasters like floods, hurricanes, wildfires, and earthquakes became the work of the EM professional. Since the early 1970s – just as the civil defense era was ending – their primary tool has been some version of the incident command system (ICS). 

ICS is a command and control system developed by and for the military and adopted by first-responder groups like fire and police departments. Its primary goal is to combine the management of a disaster response into a coherent leadership group answering to the same chain of command. Working out of an emergency operations center (EOC), the response group leaders work together under one command to support what might be dozens or hundreds of responders on the ground.

BCM

Business continuity management (BCM) came about at the same time, but for different reasons. The peak of the B-52 nuclear bomber age was also the start of the computer age. Mainframe computing came into popular use in the late 1950s and grew exponentially from its start. Major industries came to rely on their mainframe “big iron” to keep their books and process their transactions. Activities that used to take hundreds or even thousands of employees could now be done in a fraction of the time with little or no human error. The IT adage “garbage-in-garbage-out” has always been true (bad information going into the system will produce bad information coming out of the system).

By the 1960s, companies realized their computerized data was their livelihood. The loss of data would mean significant financial and reputational losses. It would mean the loss of customer data and transactional data. If a major airline lost its reservations data, for example, the results would be chaotic. 

From the need to safeguard computer data there came a new industry – disaster recovery (DR) computing. DR centers were created to copy data and other computing assets from production data centres to backup centres where it could be recovered. The increasing immediacy of the need for this recovery led to more and more investment in DR capabilities. Imagine how quickly that major airline would need its reservation system back after it had failed. 

Companies needed not just their data but their entire IT systems back when they crashed – right away. The IT focus stayed with the DR industry through the 1980s. In the 1990s managers started to recognize that without the recovery of the workforce and a means for employees to work, the recovery of IT systems and data might be irrelevant. By the time the Y2K Bug was an issue (just before the year 2000), workplace recovery plans had started as an extension of IT/DR plans. 

The IT/DR roots of BCM are with us to this day. Richard L. Arnold founded Disaster Recovery Journal and a similarly-named certification organization in 1987. Both organizations are rooted in the original technological approach to BCM via the recovery of computer systems. Because they originated in the technology support of business processes, BCM has focused on the business functions of an organization. In most cases, IT departments serve as a support for businesses processes. For example, a major bank may have $100 million in computing assets, but their real business is still banking, not technology. 

BCM programs formed from the bottom up. Most plans are done at the operational level, planning for the recovery of key value-added tasks done by individuals and teams of employees. Separate crisis management plans (CMP) were developed for organizational leadership to make critical business decisions. Those CM groups have a similar function to the EM practitioners and their formal EOCs. The strategic decision-making is communicated to the tactical and operational recovery teams.

The difference in the BCM approach is that the planning and execution happens at the grass roots of the organization, not at the top. Since IT technicians are responsible for technology, support for the organization at the most basic level. They are responsible for each keystroke that enters information or produces a transaction. Their perspective, therefore, tends to be bottom-up. BCM planning, as a next step from IT/DR planning, also starts at the base of the organization.

BCM organizations usually have a crisis management team (CMT) consisting of senior leaders in the organization. These teams function as a less formal EM type of team. Senior leadership in BCM organizations are not usually well trained in emergency response – that’s not their job. 

CMTs usually have soft plans that largely mirror the organizational structure of the company. They trigger the activation of BCM plans, often because that activation incurs costs that only they are authorized to spend. After an activation, their main job is to communicate internally and externally and to monitor the overall response. They do not train heavily on crisis response because the time it takes from their schedule is usually not deemed valuable enough for the heavy cost in executive hours.

BCM programs are built up from business impact analyses (BIA) that is done at the business unit level. They are sometimes combined at higher levels of the organization to create tactical and strategic BIAs and plans for more senior leadership. Often this step is not taken, and BCM planning remains at the operational level.

Different Focus

Coming from emergency management, where the main focus is to protect a community and support emergency response, COOP professionals are outward-looking in their approach. Their job is to reach out and help others. They protect, heal, and put fires out. Emergency managers will lead an EOC – large or small, formal or informal, depending on the circumstances of the emergency. Their job is to effectively marshal available resources to address the emergency at hand. They work in the room with the maps and screens and people wearing a variety of uniforms. COOP is delivered using the tools emergency managers know well – the senior level EOC directing the organization’s response from the highest level.

Coming from the foundation of the business, BCM is inward-looking. The focus is on the preservation of the organization – staying in business. The first part of any BCM program is to “understand the organization.” BCM programs map the processes of each business, show how processes intersect, and how they are dependent on other internal processes. Impacts on the organization from the loss of the process are carefully measured. Resources needed for recovery are listed. Because of the IT origins of the profession, BCM begins in the details, including which computer applications are used by which business processes and how recoverable the data must be. 

The BCM perspective is therefore business process oriented. BCM professionals are often familiar with IT systems and work with business units to understand IT recovery needs based on the business process drivers. DR recovery times – the time it takes to restore systems after an outage – are dictated by the needs of the business. BCM planners are often the go-betweens communicating recovery requirements from the business to IT. DR testing supports business requested recovery requirements from the BIA, and BCM governance helps ensure those requirements are met in DR testing.

COOP organizations are not as linked to their IT supports. Because computing in the public sector has not had the same drivers of cost and profit as the private sector, their emergency plans have not had the same focus on recoverability. The plans are done from the strategic level down, and they only capture IT recovery needs at a macro level and do not usually work to the micro level of the business process.

Recovery

BCM and EM were both born, in part, from the California forest fires of the early 1970s. The ICS system was formalized in that crucible by the National Fire Prevention Association (NFPA). So was the NFPA-1600 standard, the first standard dedicated solely to BCM. 

For both approaches, the plans must encompass recovery over time, through either emergency management or crisis management response. These are short-term activities that cover the initial strike of a disruption. After that first response, the COOP or BCP program kicks in to keep the organization going over many days, weeks, or even months. If a disruptive event is over quickly and the organization can get back to work in a short time, neither COOP nor BCM plans are necessary. 

Emergency management has embraced the challenge of keeping other parts of their organization working during a disruption by their COOP programs. EM responds to the cause of disruption (the hurricane, food, or fire). Their view has been expanded via the COOP programs to recover other parts of the EM organization in ways that seek to accomplish the recovery goals similar to BCM. Responder organizations (police, EMS, fire services, etc.), governments at various levels, non-governmental organizations (NGOs), and other services recognize the need to keep themselves operational. Governments need to stay up and running to provide the many services on which their citizens rely. Utilities need to keep checking meters, running their offices, and doing regular operational maintenance in places outside the disruption. 

To that end, COOP organizations choose critical business operations from a top-down set of criteria. Decisions about the order of recovery and the time-criticality of functions are made from above by more senior management. Criticality and recovery time are joined as one concept. If something needs to be recovered sooner, it is more critical. 

The same approach to criticality and recovery time is taken by IT. In the technology world, recovery priorities are set by determining which business units need to recover first. An earlier recovery time means, necessarily, a more important business process.

In BCM, recovery time is set by determining impacts. But the two measures – recovery time and business process importance – are two separate categories. Some business processes can have a higher criticality with a longer recovery time. A business process can be vital to the existence of the organization but have a longer time needed to recover. It is more important to BCM organizations to identify the more critical processes regardless of time, because they might not recover every process. Some will be left by the wayside until the crisis is over. Some processes will need to be recovered before a fixed deadline – later – but are absolutely necessary to the functioning of the organization. 

EM involves incidents that could have a disastrous impact on life and property. COOP response focuses on supporting first responders and others who are dealing with the disaster directly. COOP plans support responses in the field and challenges faced by the loss of facility, personnel, or IT. Because of the importance of the EM role, the COOP role can sometimes take a lesser role in the organization.

In the effort to cross the line from emergency response to continuity of business operations, COOP planning takes the same approach to continuity that it does to emergency response. That is the tool set which works well for the EM response. Why would it not be right for COOP? Because there are few or no limits to recovery resources in a COOP organization, a top-down approach to organizational recovery may seem appropriate. 

Spare No Expense

COOP engages in planning for organizational continuance for what they deem critical services. These services can be finance, HR, or other “normal” organization pursuits. They can also involve the continuation of life-critical services such as hospital operations, city works operations, drinking water processing, and other services that a community cannot do without. These services need to continue over time and will be restored no matter the expense. 

Most lifesaving organizations use COOP as both a support for emergency response and to keep their more prosaic business operations up and running. COOP supports both lifesaving emergency operations and organizational continuity. 

BCM focuses on the survival of the of the organization. The crisis management aspects of a BCM program focus on the strategic decision-making by leadership to manage the organization. The purpose is to save the organization and try to preserve its value. 

COOP plans are appropriate for business units that have a low, near-zero recovery time objective (RTO) and a high criticality. Whether the organization is private or public, non-profit or for-profit, if the service or business process has no room for downtime and is critically important to the continuing operation of the business, it must be treated as a COOP plan. Other parts of the organization that are less critical or have a longer recovery time – and can safely endure downtime – should be planned for with a BCM approach. 

First responder groups are always COOP planned. They have no tolerance for downtime and are critical to the life safety of the community. In a private organization, groups like call centers require COOP plans. Call centers are often 24-7 operations and represent the reputational and financial backbone of many companies. Not answering the many calls they receive can have dire effects on the business. Their continuity plans must recognize this and be capable of immediate and long-term recovery.

Perhaps the biggest difference between COOP and BCM is the ability of the organization to survive, and the capability of the organization to recover from a disaster. COOP organizations expect they will do whatever is needed to meet the challenges of a disruption. There is no question they will remain operational during and after the event. 

BCM organizations know they can fail. They seek to remain viable as an entity, with the understanding an unsuccessful recovery may force them to go under. 

As a function of government or a utility or other permanent organization, COOP planners have no practical limits to their response to a disaster. If more resources are needed to address a flood or fire or even an IT/DR event, those resources will be brought in. Even a small municipality can declare an emergency and bring in resources from increasingly higher levels of government to address a disaster. The national governments will expend resources and even bring in other governments, NGOs, or UN responses when they are overwhelmed. 

If a municipality is fighting a forest wildfire, for example, they will not stop when their fire department is overwhelmed. They will reach out to the state or province for more firefighters, more equipment – whatever they need. In the 2017 fire response in British Colombia, firefighters from other provinces, the US, and Mexico were dispatched to the fire. 

BCM has a different job: the survival of the organization with limited resources. The scope of its recovery is limited by the cost-benefit analysis they must consider. Is the business segment worth what it will cost to recover it? 

Any city, region, state, province, or country will still be there after a disaster. The city of Mississauga will go on, even after a major disruption and even after a significant loss of life and property. The same cannot be said for BCM organizations. Organizations can, and do, cease to exist after major disruptions. At some point, a BCM organization can come to the end of its ability to cope with the disruption. At that point, they cease to exist as they were and either reduce their operations significantly or simply end operations for good.

Surge Capacity

COOP organizations are also concerned with the surge of resource requirements. As a disaster event escalates, COOP organizations continue to draw in more and more resources. Some parts of a COOP organization are reduced by the loss of personnel who are assigned to long placements in disaster response. They need to be replaced in their usual roles by a surge in resources brought in from other places or organizations. Other parts of the organization see a surge in their workload they must address. For example, a disaster in a residential area would cause a surge in lost animals. Municipal animal services would have to take in and care for many pets until they can be returned to their owners. 

BCM organizations are concerned primarily with their reputation and brand value. There is no amount of money that can buy back a tarnished brand. If there is a public perception the organization has failed to do its job and has put people in harm’s way or has otherwise acted malevolently, that damage is nearly impossible to reverse. COOP is also concerned with reputation, especially if their agency is closely tied with a politically elected or appointed leader. Politicians can lose their jobs if an EM or a COOP response is not perceived as being properly handled. 

Secondarily, BCM organizations watch their money. That is not necessarily because they are trying to preserve profitability. Companies are aware that disasters are expensive, and most can rely on an insurance payout down the road. A disaster which rapidly drains money and the ability to make more money can be fatal to the organization. Most companies do not have unlimited credit and do not have the ability to draw on another organization to stay afloat the way a COOP organization can rely on higher levels of government. Also, public companies rely on their share price in the marketplace. If that price falls due to a mismanaged disaster response, the company can be severely impacted or it could fail. 

COOP leadership works to resolve disasters in their areas of responsibility using their resources wisely, but with the knowledge there are always more capabilities into which to tap. They worry about the life safety of their employees and the public. They want to offer services as soon as possible after the event. They try to preserve property and the environment. Their concern is the delivery of services which are often crucial to the life and health of a public constituency. If public transit buses cannot drive, people are left stranded on the roads. 

Utilities are also COOP planners. When a utility goes down it will expend unlimited resources to bring service back. Utilities use mutual aid agreements with other utilities. When they are overwhelmed, they can also count on government support. Even in the financially bankrupt, hurricane-devastated island of Puerto Rico, work will continue until electricity is restored. 

BCM planning is not involved as much with life safety issues. They assume first responders will do their jobs, and BCM organizations are not normally qualified to save lives. A BCM program is in place to prioritize the recovery of the organization in a way to preserve as much value as possible. In a BCM organization, this value is not normally associated with life safety.

Practical Effect

Binder tableWhen a new program is needed, a planner needs to decide whether it will be a COOP plan or a BCM plan. The best way to make this distinction is through a valid business impact analysis (BIA). When executed properly, a BIA can determine the recovery time and criticality of a business process. If a process is highly critical – if its loss would have a devastating impact on the organization and has a “zero recovery time” – a COOP approach is needed. The BIA can show that a process is less critical or has a longer time to recover. In those cases, a BCM approach is needed.

BCM organizations are focused on the recovery of their own service and value-added activities. The program is developed from the ground level up. Senior management decides on the structure of plans, normally along the lines of the organizational chart. They are engaged in the approval of program elements when complete, and for supervision and handling of the program as a whole. 

A COOP program starts with senior management deciding the criteria for business unit criticality, and often picking the critical units for the program themselves. The COOP program views their critical services as a series for activities that must be provided to their client groups within a defined timeframe. 

Taking a COOP approach is less effective for business processes in its response to an outage. The ground-to-ceiling approach of a BCM organization is stronger in understanding all aspects of the company and the value-added proposition of each segment. The COOP top-down approach can miss many of the important details a BCM response will take into consideration.

The same problem faces a BCM approach to an EM challenge. The EM planner using the bottom-up approach would have to run the BCM program from the operational level up to the strategic level to plan for disaster response. That would entail too much detail and coordination for an EM group which faces regular disruptive events. It is easy to see why emergency managers would not need to do a process-level program for first responders.

The best approach for a COOP organization is to split its planning between emergency operations and business operations. For those parts of their organizations which do not have to respond to disasters directly, they should take a BCM approach from the ground up. That way their BCM plans will be validated to the grassroots level, while their COOP operations will not be overwhelmed by the level of detail captured in a BCM program.

In both organizations, time and effort are split between crisis or incident responses and business program responses. The crisis or incident management teams use a response structure involving strategic leadership of the organization. In BCM organizations, this is a relatively small part of their response activities. Most BCM work takes place with the business unit itself and the planners within the business who complete the BCM tasks. 

In COOP organizations, the intent is to maintain service to the customer outside the organization. Most of the work goes into preparing for and running the IMT response using the approved structure (IMS, ICS, and more). Efforts to plan for the recovery and maintenance of the rest of the organization, outside of the direct response to the disruption, are planned using a similar approach, but would be well served by using more BCM methodology. Many COOP planners are certified in business continuity management through the Business Continuity Institute (BCI) and other certification organizations. That training lends them a basic skillset in BCM, but for a deep understanding of the profession it is crucial to have experience with an organization with a robust BCM program. 

COOP and BCM planners need to make the distinction between COOP and BCM planning approaches. BCM organizations often need a more robust CMT to handle the onset of an unexpected outage. Their CMT plans should be modeled after emergency management plans as written and executed by EM groups over many years and activations. By the same token, government and other EM organizations need to take a ground-up BCM approach to their COOP programs for the aspects of their organization which do not have short RTOs with high criticality. 

In the end, knowing what your program will deliver is key. Will it be a top-down, response-heavy COOP program? Or do you need a bottom-up, detail-heavy BCM program? Almost every organization will need both in different dosages. 

Binder Abraham optAbraham Binder has a master’s degree in disaster and emergency management and a master’s degree in history. He has taught business continuity at York University in the graduate and undergraduate degree programs in disaster and emergency management. Binder has worked in private sector business continuity since the Y2K threat. He has work with major banks and insurance companies in Canada, developing business continuity programs and working with IT departments on disaster recovery. Since 2016, Binder has worked for Mississauga, Canada’s sixth largest city with a population more than 700,000, in developing business continuity and continuity of operations programs.

 

Latest News
DRJ HOT ITEMS
Webinar Spotlight
Fetching Upcoming Webinars...
Journal Categories

AI: Automation & Innovation

Business Continuity Management

Crisis Management & Emergency Response

Cyber Resilience & IT Disaster Recovery

Leadership: Culture & Workforce Resilience

Operational Resilience

Risk Management & Quantification

Sector-Specific & Critical Infrastructure Resilience

Supply Chain & Third-Party Resilience

Governance: Compliance & Regulatory Readiness

Incident Management & Response Coordination

Resilience Strategy & Program Maturity

Data Protection: Backup & Recovery

Exercises: Testing & Scenario Planning

Emerging Threats: Geopolitical & Climate Risk

Contact Us

Newsletter

The Journal, right in your inbox.