Q&A with Puneet Khatri, Head of Services, Libelle Americas

Puneet Khatri is a seasoned SAP technology leader with more than 18 years of experience architecting, managing, and securing complex enterprise SAP landscapes. As head of services at Libelle Americas, he focuses on disaster recovery, high availability, hybrid cloud strategies, and data masking for mission-critical systems. Today he speaks about how enterprise resilience is changing, where DR programs still go wrong, and what leaders should be doing differently.
Tell us a little about your background and what you focus on today?
Khatri: I’ve spent close to two decades in SAP technology leadership — Basis, disaster recovery, high availability, and data protection. Day to day, my work is about keeping mission-critical SAP environments running through disruption and constant change, across on-premises, AWS, and Azure. I manage a large client portfolio, so I see the same resilience challenges play out across very different organizations. That pattern-spotting is what shapes most of my thinking.
You’ve been in this space for nearly 20 years. What’s the biggest way disaster recovery has changed?
Khatri: The definition of “recovered” has changed. For years, DR was infrastructure-centric — a secondary site, a replication link, an annual failover test. Success meant a server came back. Today, environments change daily: cloud, automation, frequent deployments, dynamic security controls. So the question is no longer “Did the system restart?” It’s “Did the entire business service — data, identity, and dependencies — come back correctly, under real conditions?” We’ve moved from recovering infrastructure to sustaining resilience.
What’s the most common misconception you see in DR programs?
Khatri: That a passed DR test means you’re resilient. A test validates a snapshot under controlled conditions — usually off-hours, scoped to avoid the messy scenarios. Real incidents don’t schedule themselves for the weekend or avoid peak load. A passed test can manufacture false confidence, and false confidence is more dangerous than acknowledged risk.
You’ve made the provocative argument that “DR testing is dead.” What do you mean by that?
Khatri: Disaster recovery isn’t dead — the DR test, as an occasional ceremony, is. The core problem is timing. Validation has to move to change time, not disaster time. Every architecture change, security update, or access modification introduces new risk. If you only discover the impact during the annual test — or worse, during a live event — you’ve already lost. Continuous resilience validation means small, controlled, ongoing stress that surfaces fragility while it’s still cheap to fix.
Your recent work points to non-production systems as a hidden resilience risk. Why should leaders care about them?
Khatri: Because nobody calls them critical, so nobody protects them — yet they hold full copies of production data and often connect straight back into production. In a real event, the “non-critical” sandbox is exactly where a breach or a compliance failure walks in. And here’s the uncomfortable part: a DR copy is itself a non-production instance of production. If you govern those environments casually, you’ve quietly built your recovery on your weakest-controlled systems.
Data masking and security are central to your work. How do they connect to resilience?
Khatri: Resilience isn’t only “does the data come back” — it’s “does it come back safely.” If a recovery or a system refresh exposes unmasked personal data or opens an unguarded access path, you haven’t removed the risk, you’ve relocated it. Masking sensitive data at refresh time, and disciplined identity and access control, are part of resilience — not a separate security checkbox. Regulators certainly don’t treat them as separate.
Where do AI and automation fit into the future of DR?
Khatri: The direction is from reactive recovery to predictive operations — using telemetry and signals to act before failure, and policy-driven orchestration instead of human heroics under stress. AI won’t replace human judgment, and it shouldn’t. But it can shrink the window between signal and response, and it can take the fragile “manual decision at 3 a.m.” out of the critical recovery path. That’s where the real reliability gains are.
SAP is your specialty. Is there anything distinct about resilience for SAP landscapes?
Khatri: SAP isn’t one system — it’s a constellation. A production tenant surrounded by many downstream copies, tightly coupled databases, and deep dependencies. In practice, recovery often fails at the data layer, not the application layer. And every one of those copies carries sensitive data. So resilience for SAP is as much about data governance and dependency mapping as it is about failover mechanics. Teams that only rehearse the failover, and never the data and dependency paths, get surprised.
What should resilience leaders do differently, starting Monday morning?
Khatri: Three things. First, inventory every environment — including the ones you label non-critical — and classify what data each one holds. Second, move validation to change time, so risk is caught when it’s introduced. Third, validate the recovery path, not just the recovery target — the identity, access, and masking around recovery, not only whether the data returns. None of that requires a bigger budget. It requires that you stop trusting labels and start testing reality.
If you had to leave leaders with one line, what would it be?
Khatri: Resilience is demonstrated, not documented. A binder full of plans and a passed test are not the same as knowing — this week — that you can actually recover.

