drj logo
drj logo

Welcome to DRJ

Already registered user? Please login here

Create new account
(it's completely free). Subscribe

x

IT-DR Trends to Concern Business Resiliency Leaders

Cyber Resilience & IT Disaster RecoveryOperational Resilience

For the last several years, the Data Protection Trends Report has published the unbiased sentiments of IT leaders who were responsible for the backup and (IT) disaster recovery strategies for their organizations. The 2024 report surveyed 1,200 enterprises with some unsettling insights for senior leaders responsible for the business continuity or crisis management of their organizations – especially for C-level executives who may have delegated key aspects of their BC/CM or business resiliency strategies to their IT leaders.

What is causing IT outages?

Ransomware? Absolutely, but the survey data reveals if organizations have overly conflated their cyber-preparedness with their IT-DR strategy, then they are likely underprepared for the myriad crises that continue to plague organizations of all sizes. Yes, for the fourth year in a row, the survey reveals ransomware and cyberattacks were not only the most common cause of outages, but also the cause of the single most impactful outage of each of the past four years. As business resiliency planners, we are right to be ever mindful of how we ensure the cyber-resilience of our organization. However, the media, the vendor community, and many consultants/partners have over-rotated.

According to the 2024 report:

  • IT component failures including various hardware and software continue, with very little abatement year over year.
  • Outages of cloud-hosted resources is actually rising in frequency as most organizations have now embraced a “cloud-first” strategy. In fact, if you consider connectivity issues to the cloud alongside outages caused within the cloud, this concern tops most other areas.
  • Lastly, natural disasters, while not nearly as “glamorous” as cyber, continue with the seasonality many long-tenured BC/DR experts have always known about and planned for.

In other words, while cyberattacks continue to gain mindshare, none of the other kinds of crises that debilitate IT, and affect business processes, are receding. The caution to BC/CM executive stakeholders is:

Ransomware is a ‘when’ not an ‘if’

That does not mean assured destruction and utter calamity is a “when” not an “if” – just like any other crisis at scale. That said, unless natural disasters, supply chain issues, or the myriad other crises BC/CM teams have to plan for – ransomware comes with no warning … other than the constant media coverage and articles like this one reminding you of its inevitability.

According to the same 2024 research, 25% of enterprises don’t believe they suffered one attack in the preceding 12 months, while 26% acknowledge they were hit four or more times over that same time period. Said another way:

That isn’t even the worst of it when you consider cyber villains are often lurking and navigating throughout the environment for up to 200 days before they are ever discovered or announce their demands. As such, many of those who believe they haven’t been hit have been breached, but are blissfully unaware.

It is important to note that repeat attacks do not entirely imply multiple breaches. As discussed later in this article, many organizations simply fail to completely eradicate the initial malware. Months after the first cyber event is over, the attacker simply checks if any of their technology is still present in the victim’s environment.

In this case, the 2024 Ransomware Trends Report reveals only 37% of organizations have the ability to sandbox or quarantine a staged restore via a “clean room” to ensure they do not re-infect the environment. As any seasoned BC/DR professional knows, there is often a significant amount of pressure to resume operations as quickly as possible. Without the proper orchestration or planning, victims might remove malware from their systems only to inadvertently re-infect themselves during data restorations.

There is an unfortunate nuance in recovering from ransomware at scale, even if you have a mature and well-orchestrated IT-DR program. With fire or flood, the secondary data is good right up until the original systems became crispy or wet. The IT teams’ mandate is to simply re-home and enable the secondary systems as quickly as possible. Unfortunately, the secondary data from minutes before the ransomware event are likely just as infected as the production instances, so even the triage of those systems affected become a non-linear and non-predictable delay before remediation can begin in earnest.

There is also the alarming reality cyber villains target your backup repositories in 96% of attacks and are successful in 76% of the attacks to encumber or eliminate your IT teams’ ability to recover your data – increasing the likelihood you’ll have to pay the ransom.

IT teams are testing DR less

With as much hype as ransomware has in the media, as well as various regulatory mandates for cyber-resilience and board-level directives, one would assume testing of IT systems’ recovery at scale would be on the rise. Unfortunately, the 2024 statistics do not show that. Many BC/CM teams saw additional testing and plan development efforts in the years immediately after COVID, based on teams’ bandwidth and an inherent recognition of the organization’s depending on IT systems, even as some of those IT systems evolved from datacenter centric to new architectures in support of remote workforces.

Research from 2024 reveals IT teams are only testing recovery-at-scale (e.g. IT disaster recovery) every eight months, which is less than the “every six months” in 2023 or “every five months” in 2022. When IT leaders were asked what percentage of their production systems was recoverable within expected SLA’s during their last IT-DR test, only 58% of systems came online within their timeframe. While others likely simply resumed slower, one has to presume some percentage weren’t recoverable at all.

IT-DR SLAs are not meeting expectations

As resiliency professionals, we know the secret to any successful resiliency plan is in the consistency by which it can be executed (barring unforeseen circumstance variation). Surprisingly, only 37% of IT teams utilize orchestrated workflows as part of their systems’ recovery processes. Some 37% can test their “recipe” for server/application recovery, examine which workflow steps exposed errors, and then optimize the workflow for higher reliability in the future. The other 63% are conducting manual recovery steps per workload, each SLA adherence will vary by test, with little to no opportunity to improve the potential agility in future exercises or actual events.

In fact, when enterprise IT leaders were asked how long they would anticipate the IT-DR recovery of a relatively simple 50-server environment (which isn’t a lot in an enterprise environment), only 32% believed they could recover those 50 servers within a single business week.

Conclusions and recommendations

There are more statistics related to the potential for IT teams to meet the recovery expectations during “typical” disaster recoveries as well as ransomware events in the research cited above, including (data protection trends) and (ransomware trends). In the meantime, as business resiliency leaders, here are a few questions to discuss with your DR constituents in IT:

  1. How do we ensure our backups are not affected when ransomware hits our production systems?
  2. When was the last time we tested at scale (e.g. 50 servers or more)?
  3. How much of the per-workload recovery is orchestrated with workflows that can be assessed and optimized?
  4. Do we have a “clean room” or other staged restoration capability to reduce the risk of re-infection during restoration from a cyberattack?
  5. When were these IT-DR capabilities last externally audited against our expected SLAs and recovery plans?

That last one ought to make the other four much more instructive.

ABOUT THE AUTHOR

Jason Buffington

Jason Buffington has spent 35 years in IT disaster recovery. He first earned his CBCP in 2003, spoken at numerous DR and IT events over the years, and has been published in DRJournal and other periodicals. He is a VP of strategy at Veeam Software and his blogs can be found on http://ITDRblog.com. View the entire Data Protection Trends Report at http://vee.am/DPR24.

Latest News
DRJ HOT ITEMS
Webinar Spotlight
Fetching Upcoming Webinars...
Journal Categories

AI: Automation & Innovation

Business Continuity Management

Crisis Management & Emergency Response

Cyber Resilience & IT Disaster Recovery

Leadership: Culture & Workforce Resilience

Operational Resilience

Risk Management & Quantification

Sector-Specific & Critical Infrastructure Resilience

Supply Chain & Third-Party Resilience

Governance: Compliance & Regulatory Readiness

Incident Management & Response Coordination

Resilience Strategy & Program Maturity

Data Protection: Backup & Recovery

Exercises: Testing & Scenario Planning

Emerging Threats: Geopolitical & Climate Risk

Contact Us

Newsletter

The Journal, right in your inbox.