For years, cloud migration has been associated with greater resilience. Moving workloads from on-premises infrastructure into hyperscale cloud platforms promised better availability, geographic redundancy, and less dependence on aging hardware. In many ways, that promise has been fulfilled, giving organizations greater flexibility and scalability than they could achieve on their own.
At the same time, experience has shown cloud providers and the services surrounding them are not immune to outages. These disruptions are not always caused by cyberattacks or natural disasters. They can result from configuration errors, software issues, networking problems, or failures within shared services that ripple across thousands of customers.
The complexity of modern IT environments has made these events increasingly difficult to avoid. Uptime Institute’s 2025 Annual Outage Analysis discovered IT and networking issues accounted for nearly one-quarter of significant outages in 2024, while third-party providers (including cloud, telecommunications, and internet services) have been involved in roughly two-thirds of publicly reported outages over the past nine years. The same report found more than half of organizations experiencing a significant outage estimated the cost exceeded $100,000.
The question for IT leaders is no longer whether an outage will occur. It is whether the business can continue operating when one does.
Resilience Is More Than Disaster Recovery
Many organizations still equate disaster recovery with business resilience, but the two serve different purposes.
Disaster recovery focuses on restoring infrastructure after an incident. Business resilience focuses on maintaining operations while an incident is still unfolding. The objective is not simply to recover systems as quickly as possible, but to minimize disruption to employees, customers, and the business itself.
That distinction has become increasingly important because today’s enterprise environments depend on a growing number of external services. Cloud infrastructure, identity providers, SaaS applications, storage platforms, networking services, and collaboration tools all represent dependencies organizations do not fully control. A failure in any one of those services can have consequences well beyond the original outage.
In many cases, the organization’s own infrastructure isn’t what fails. The disruption occurs because one of the services it relies on becomes unavailable.
The Risk of Concentration
Standardizing on a single cloud platform simplifies operations, but it also creates concentration risk. When critical applications, user authentication, desktop environments, and productivity services all depend on the same provider, a single outage can affect a large portion of the business.
This is becoming a growing concern for enterprise risk teams and regulators. The European Union’s Digital Operational Resilience Act (DORA), for example, specifically addresses ICT concentration risk and the growing dependence on critical third-party technology providers. The regulation reflects a broader shift toward ensuring organizations can continue operating when key technology services become unavailable.
The conversation has therefore expanded beyond high availability. Organizations are increasingly asking whether they have too many critical services tied to a single provider, identity platform, or cloud ecosystem.
Multi-Cloud as a Resiliency Strategy
For many years, multi-cloud discussions focused on selecting the best capabilities from different providers. Today, the stronger business case is resiliency.
That does not mean every workload belongs in multiple clouds. Replicating everything would introduce potentially significant cost and complexity. Instead, organizations should identify the applications and data sets whose failure would create the greatest business impact.
For Tier 0 and Tier 1 workloads, including healthcare systems, financial applications, contact centers, and other mission-critical services, adopting a multi-cloud approach to compute, for example, can reduce the likelihood a single provider outage disrupts the entire business, thereby reducing or even eliminating the blast radius of the outage (depending on the architecture). Less critical workloads can often be distributed across providers to reduce the overall impact of an outage without requiring a fully redundant architecture and effectively minimize the blast radius.
The goal is not to eliminate failures. It is limiting their business impact.
End-User Computing Is a Logical Starting Point
Virtual desktop infrastructure and digital workspace platforms are often among the first places organizations evaluate a multi-cloud strategy because employee productivity depends on them. Modern desktop delivery platforms can separate the user experience from the underlying infrastructure, allowing desktops and applications to be delivered from multiple cloud environments while maintaining a consistent access experience.
That flexibility becomes valuable during an outage of a provider’s compute services in particular. Rather than rebuilding infrastructure after users lose access, organizations can continue delivering desktops from an alternate environment that was already part of the design.
The most resilient environments move beyond thinking about disaster recovery as an emergency event. Instead, they are designed so multiple environments are available, regularly tested, and capable of supporting production workloads as part of normal operations.
Resilience Requires Looking Beyond Infrastructure
Deploying workloads across multiple clouds does not automatically create resilience. Organizations must evaluate every dependency that supports those workloads.
Identity is one of the most common examples. A company may successfully deploy desktops across Azure and AWS while relying entirely on a single identity provider. If that provider becomes unavailable, users may be unable to access either environment. Adopting multiple identity providers has its own challenges however, and organizations need to determine just to what degree and what failure scenarios they are willing to mitigate for.
The same principle applies to application back ends, databases, profile services, file storage, and networking. A resilient desktop environment offers little value if the applications users depend on remain unavailable because another component has failed.
Supporting multiple cloud providers also introduces additional operational complexity. Different platforms have unique networking models, provisioning methods, security controls, and management processes. That complexity should not be underestimated, but for organizations supporting mission-critical operations, it is often far less costly than the impact of a prolonged outage.
Turning Planning Into Practice
One of the most effective ways to prepare for outages is through regular tabletop exercises that involve business leaders, application owners, infrastructure teams, security professionals, and operations staff.
These exercises should examine practical scenarios. What happens if a cloud region becomes unavailable? What if the identity provider fails? Which applications must be restored first? Can employees continue serving customers while key systems are unavailable?
These conversations frequently uncover hidden dependencies that never appear in architecture diagrams or disaster recovery documentation. Just as importantly, they help organizations determine which systems truly justify additional investment in resiliency.
Designing for Business Continuity
Cloud outages are no longer rare events. They are an expected part of operating modern digital infrastructure. Hyperscaler providers will work to improve SLAs and minimize unplanned outages through experience, but other surprise outage events are still likely to crop up. While organizations cannot prevent every provider disruption, they can reduce the extent to which those disruptions affect the business. That requires moving beyond disaster recovery and designing architectures that assume failures will occur.
The organizations best prepared for the future will not be those that simply recover the fastest. They will be those that have deliberately reduced single points of failure, diversified critical dependencies where appropriate, and built environments that allow the business to continue operating even when part of the technology stack becomes unavailable.


DOWNLOAD EXCEL
DOWNLOAD WORD DOC
DOWNLOAD PDF OF EXCEL 



