Home›Insights›Articles›Disaster recovery for SAP on IBM i: lessons from a real failover

Blog · Business continuity · Retail

Disaster recovery for SAP on IBM i: lessons from a real failover

Every contingency plan looks good on paper. What we learned over more than three years protecting 80 TB of SAP on IBM i across two cities, including a real failover in July 2026.

Why continuity matters so much in retail

Colombian retail is large and concentrated: the 20 biggest companies posted combined revenue of COP 117.3 trillion in 2025, up 12.3%, and the top five captured 65.2% of that revenue, according to El Colombiano. At that scale, every hour of downtime is measured in lost sales, purchase orders that never reach suppliers and shoppers who walk over to a competitor.

The chain in this story runs about 420 stores, roughly 4,800 point-of-sale terminals and close to 8,000 SAP users, plus more than 200 SQL Server databases around the core. An outage at its primary data center isn't an IT problem; it's a whole-business problem.

RTO and RPO: the two numbers that matter

Every recovery strategy boils down to two targets the business must set and technology must meet:

  • RTO (recovery time objective): how long the system can be down before it's running again. Here, 3 hours for the SAP core after a major outage.
  • RPO (recovery point objective): how much data can be lost, measured in time. Here, a maximum window of 15 minutes in replication.

With 80 TB of data, those targets are very hard to hit with traditional backups alone: restoring a database that size from periodic copies typically takes far longer than a few hours, and you can lose everything since the last copy. The only realistic way to meet them is to keep a live replica of the environment at another site.

How IBM PowerHA for IBM i works

IBM PowerHA is IBM's high availability and disaster recovery solution for the Power platform. On IBM i it integrates natively with the operating system and provides:

  • Cross-site clusters that group the primary and secondary systems into a single solution.
  • Continuous data replication , either at the operating-system level or built on storage-based replication, depending on the design.
  • Controlled switchover for planned tests and failover for real incidents.
  • An administrative domain that keeps user profiles, configurations and system objects in sync across nodes, so the secondary site is ready to take over.

In this project, the 80 TB SAP environment is replicated between two cities, and Redsis manages the IBM Power infrastructure at both data centers.

A recovery plan that has never been executed is a hypothesis. One that is tested continuously is a business capability.
Infrastructure team, Redsis

Five lessons from more than three years in production

1. Let the business set RTO and RPO

Recovery targets aren't a technical setting; they're a business decision about how much time and data you can afford to lose. When operations sets them, the architecture is designed to meet them, not the other way around.

2. Distance is part of the design

A secondary site in the same city protects against hardware failure, but not against a regional event. Putting the replica in another city broadens protection, at the cost of carefully designing links, latency and replication mode.

3. Test, test and test again

The solution in this case is tested on an ongoing basis, and a real failover was executed in July 2026. That's the difference between believing the plan works and knowing it does. A good practice is to alternate planned tests with exercises that come as close as possible to a real incident.

4. Look beyond the core

SAP is the heart, but it isn't alone: point-of-sale systems, integrations and satellite databases also have to come back. Documenting the recovery order for the whole ecosystem prevents a situation where the core is up but the stores still can't sell.

5. Run high availability as a service

Replication degrades when nobody watches it: configuration changes, data growth or upgrades can quietly break it. Having a team that manages the infrastructure at both sites, as Redsis does in this case, keeps the solution ready for the day it's needed.

The result: continuity proven, not promised

Today the chain's SAP core has a 3-hour recovery time and a maximum data-loss window of 15 minutes, on a solution that has been in production for more than three years and proved its worth in a real failover in July 2026. For an operation with 420 stores and 4,800 point-of-sale terminals, that means something very concrete: if a major outage hits, the business knows how long it will take to come back and how much it could lose.

Where to start

If your organization doesn't have clear RTO and RPO targets, or your contingency plan hasn't been tested in the past year, a continuity assessment is a good first step. At Redsis we combine more than 25 years of mission-critical infrastructure experience in Latin America with our IBM Platinum partnership to design, implement and run high availability on IBM Power and IBM i.

Read the full story

See how one of Colombia's largest retailers recovers its SAP core in 3 hours with IBM PowerHA.

View success storyTalk to a specialist