The resounding success of more than a hundred BCP/DR drills that I have done with my clients actually worries me! In delivering business continuity (BC) and disaster recovery (DR) services over the years, I have closely worked with clients in making every single exercise a success. Those are our moments of glory, the BC/DR managers are in limelight, and it is important to partners in building credible references. The partnerships generally work well, and it is good business.
But deep down in my heart I reckon that even while we are preparing for a real disaster, we’re not being “real” in our preparations for them. There are many reasons, foremost being the compulsions of the BC/DR profession itself. To this community, a successful exercise is a question of credibility; it justifies the investments the company has made, of which we are the custodians. The tests reiterate that the BC/DR plan is still relevant and works. To partners and service providers, making exercises successful is as important. The consequences of a failure can be detrimental to the relationship. But is our obsession with making every exercise a success, a real test of BC/DR preparedness, or is there a fundamental disconnect? Ask yourself these questions to know why:
Will a real disaster allow you the same level of planning as the BC/DR exercises?
The preparation for every single exercise I have executed with clients started weeks if not months in advance. Micro-planning down to the last detail is the norm, and a BC/DR professional is measured by the ability to identify and plan the details. The focus is on ensuring not a single glitch in any process – no delays, no interruptions set on the premise of seamless communications, and a flawless plan that ensures the right people at the right time executing just the right steps to recover the systems and processes. We come out smiling from these wonderfully executed exercises and consider the job well done. Often, there is no follow-up.
Contrast this with a real disaster such as a fire in the building, a power grid failure, floods, or a cyclone that paralyzes basic communications and transport systems around you. Would we still have such a wonderful symphony of actors performing impeccably with every team member available at the appropriate time and place? Would an anxious and thin management, already under stress, be responsive to mobilize the support functions? For those of us who have experienced such situations, we know only too well the reality is almost always very different.
The tolerance of BC/DR community to confronting potential failures is very low. So even while we realize the importance of leaving some variables unplanned, we are skeptical about testing them. The need is to get closer to reality and not vie for perfections. Take your management into confidence and instead, introduce imperfections and challenge the situations you test (e.g. ask key people they cannot participate as they may be impacted by the disaster and see how the stand-ins perform). Get the jigsaw puzzle pieced together anew in every exercise and learn from it. Ensure every exercise has an after-action report – to evaluate what did and did not work. Capture the lessons learned and bring about the changes. A BC/DR manager who does not recall significant learnings from the last few exercises is probably missing some very important lessons.
In planning the test, did you actually test the plan?
Let us try and understand this with a few commonly seen examples. Some clients who have made investments on the basis of an earlier business impact analysis and have IT/DR solutions that enable the production applications to failover to an alternate site in the eventuality of an outage. Sadly during drills, complete applications failover is almost never tested. The tests are confined to one or two applications at a time. I have known many clients who test either the core ERP application or the mail, but never both together and some applications are never really tested for recovery from the DR site. The nuances of testing one application and calling the exercise a success are very different from testing half a dozen applications or several times that – against a ticking clock – to achieve the recovery point or time objectives committed to the business.
The management, though mindful of the investments made, generally does not understand this crucial difference. These quick wins also keep the BC/DR manager’s jobs going, the management happy, and the auditors quiet. But wouldn’t reality be different? What would be the impact on business if the enterprise mails stopped for a few days or weeks? Would recovering the ERP application only, work for your customers and suppliers, and still protect your company’s business? Do not overlook the need to test resource planning during exercises. Unless that is done, the most daunting of the challenges (i.e. enablement of administrators to work effectively and the sequence of activities in switching multiple applications) remains untested.
Citing “business risk” as the reason for not undertaking large-scale tests will leave you vulnerable. Yes, such tests will disrupt usual business for short durations, but be aware that the cost of not testing real situations can be far higher.
Are your BC/DR plans and artifacts ready for an impromptu audit?
Frantic activity precedes audits. In those few days, the plans are updated, DR calendars plugged, and historical MIS consolidated. Sometimes change management is done retrospectively to satisfy the auditors to ensure changes made at production are applied at DR. The “schoolboy tactics” to complete unfinished assignments at the deadline lead us to expired passwords, obsolete documents, outdated SOPs, incomplete and untested call trees, and similar broken threads stumbling out of the cupboard at the eleventh hour.
We accept these as professional hazard but to be fair to the profession, the realization that BC/DR plans are not for the auditors but for ourselves is foremost. Unfortunately, risks associated with poor BC/DR preparedness are too low on the management’s priority list round the year. These continue to remain so, either until an auditor comes by or a real disaster strikes to shake us out of the complacency.
It is not difficult to set up a robust governance mechanism to regularly revisit your BC/DR preparedness, without waiting for the auditors. Even if discomforting, push discussions around BC/DR ROI and roadmap on the board’s agenda. It will only help to keep the management continuously involved.
Does your business accept downtime for testing DR as an essential investment?
CXOs understand business continuity and disaster recovery from the perspective of money invested but rarely the time and effort needed to make it work. For instance, IT/DR, largely driven by head IT, often fails to take key stakeholders from the business along. This, in my observation, is because the bridge between the business and IT departments is often weak.
IT departments traditionally organized under infrastructure and application groups are more inward looking than outward. Often, none other than the head IT/CTO really understands the language of the business. Senior IT executives, while rich in technical skills, may lack the qualifications or ability to connect with the business and justify and defend the time and effort required for testing the DR. The common objectives that such investments make the business resilient and will protect the company’s reputation are often lost. Not surprisingly then, “down-time” traditionally remains a bad word for the business folks.
I have known clients with elaborate DR solutions in place, but lacking in will power to test the scenarios because the test windows would appear as down times for some other applications, which are either not protected or are too cumbersome to test. A real disaster, however, will be interspersed with down times. Assuming you already have a functional DR for critical applications, those will be brought up. But then there will obviously be those that are not protected, and your business will need to survive without them too.
Get your management to confront the fact that flashing a “planned downtime” message to real clients and accepting its impact during business as usual (BAU) is important. Each well spent downtime is actually also an investment into making the systems more disaster proof. Initiate these conversations and elevate them because that is the very objective of the investments your company has made in BC/DR, in the first place.
Have your BCP/DR strategies kept pace with dynamic baselines?
Investments into BC/DR are generally made based on 1) a conscious management decision based on business impact analysis, 2) when the company or competitors or businesses in proximity face a real disaster, or 3) to meet regulatory requirements. While the decisions are taken against a set of predefined needs and objectives that remain unchanged for many years, businesses in reality change every week. Databases grow, transaction levels increase, branches are added, new applications are introduced and sometimes even the core business processes undergo significant changes in short spans of time.
BC/DR managers are generally averse to constantly changing baselines – it means repeatedly preparing and re-preparing investment cases for the management to approve. Additionally, there could be contractual lock-ins with partners that prevent constant fine-tuning to allow most effective utilization of the created assets. When facing the management, these are uncomfortable questions.
The easiest route is therefore to accommodate the changing baselines by trading off one element with the other. For instance, network bandwidths for replication to DR are not resized for years even while the traffic grows linearly with the volume of business transactions. The list of applications protected at the DR keeps getting longer without the corresponding changes to all infrastructure elements. Technologies like virtualization make it easy to hide the underlying layers. As a result, what exists at the DR site is basically a poor imitation of the production environment. This is a dangerous mismatch as the DR site will only withstand lower performance levels, that neither the IT nor the business knows. In some cases the DR site design is consciously optimized from the start, but little thought is extended to how in reality access at the end user or application level will be controlled post disaster. The fact is that storages with reduced number of disks or servers with lower CPUs will either perform poorly or fail altogether when subjected to full load. You cannot change that. Only you don’t know it yet.
In a nutshell, while the DR site is often designed for lower loads, the rationale for such downsizing is almost never tracked and therefore lost over a period of time. In all likelihood, your management will not recall the risks that may have been accepted while the decisions were taken. The onus to keep the discussion alive is on us, but do not wait for a real disaster to get your management’s attention. By then it may be too late.
Have you thought enough about the people who run the technology?
A well-connected ecosystem fueled by Internet, social media, and competing partners ensures that you are kept abreast with the cutting edge of technology. It leads us to believe that answers to most complex DR/BCM problems can be found there. So while there is tremendous focus on technology as an enabler, the irrefutable fact that people drive technology is often overlooked.
Once, a client faced with a fire in their main office flew down 150 highly-trained agents to another city, where the BC infrastructure existed. It was all as per the business continuity plan, except that within days, anxious employees began to question how long the situation was likely to continue and when they would return. Instances of family and personal problems reported increased manifold and with the HR unable to provide clear answers, the BC plan could not be sustained. The management was forced to provide an alternate work area within the city at a very short notice but at higher costs.
I attribute poor consideration about the people to a general lack of BC/DR culture in the organization. Poor awareness about BC/DR at all levels also means that employees are not aware of their roles in a disaster situation. Even in organizations which have BC/DR infrastructure in place, it is not unusual to find key people so preoccupied with BAU routines that they see these responsibilities as unnecessary overheads. Additionally, a major disaster can also sully the company’s reputation – so much so that large-scale attrition can become a sudden reality, and the recovery process will need to run without some of the seasoned actors in it.
So what can BC/DR managers do?
My objective is to provoke you to think deeper through the questions raised above. Though there is no one-stop solution to fix the issues, sustained efforts can make a difference, provided one starts the journey. Assess your risk exposures and recovery capabilities continuously. If the maturity level of BC/DR practice in your organization is low, initiate conversations on the ensuing risks without waiting any further. However, if it is already recognized as an important function, I still recommend getting a better picture of the risk issues and controlling them through a roadmap that is relevant to the current and emerging trends in your business. Taking note of the questions asked in the previous sections may help.
For the uninitiated, a good start point could be an assessment of your company’s BC/DR posture using readily available tools. For instance, index tools may reveal how security and business continuity can shape the reputation and value of your company as against the industry benchmarks. Business continuity index tools reveal key indicators of business continuity exposure. The tool can assist in delivering better risk management strategies and raising the profile of business continuity management in your company. The index allows you to take the pulse of your organization – identifying where improvements can be made, and outlining potential next steps for your business.
Awareness is the first step. You can then choose an approach that suits you right and keep the controls. Reach out to partners who can help. Several independent studies can help you identify a broad range of successful partnerships in the business continuity and resiliency domain.
In my view, the most important consideration is what we started with: the need to get real about BC/DR preparedness. Believe me, the journey from misplaced perceptions to stark reality is an arduous and continuous one. It never stops. And just like in real life, there are no short cuts here either.
Rohit Chaudhary, PMP, is an IT professional with more than 20 years of experience and is a hands-on BC/DR practitioner. Until recently, he led the business continuity and resiliency services delivery at IBM India. He currently heads the IT project management’s center of excellence at a Fortune 200 company at Mumbai, India.
