The Recovery Testing Gap MSPs Can't Ignore in 2026
Rodney Hall, COO

Most MSPs are not testing backup and disaster recovery often enough to back up what they tell clients about recovery time. Survey data from 2025 shows fewer than one in six organizations test backups daily, and barely one in ten test full disaster recovery that often. The gap between assumed and proven recovery is the real BCDR risk heading into 2026.
How Often Are MSPs Actually Testing Recovery?
Not nearly as often as clients assume. The 2025 State of BCDR Report from Datto, built on responses from more than 3,000 IT professionals worldwide, found that only 15 percent of organizations test backups daily and 25 percent test weekly. Disaster recovery testing lags further behind that, with just 11 percent testing daily, 20 percent weekly, and 23 percent monthly.
Add that up and well over half of the organizations in that survey go a month or longer between DR tests, if they test at all. For an MSP, that number is not abstract. It describes what's likely happening inside client environments you're responsible for right now, and inside your own backup stack, unless testing is built into a fixed schedule instead of a task that gets pushed whenever the queue gets quiet.
The reason testing slips is rarely a lack of tooling. It's staffing and time. A full recovery test takes a technician off billable work for hours, and when that technician is also handling tickets, onboarding, and patch cycles, DR testing loses that argument every time unless someone above them makes it a scheduled line item rather than a best effort.
The Gap Between What You Promise and What You Can Prove
Clients assume recovery is fast because that's how backup gets sold. The same Datto research found more than 60 percent of organizations believed they could recover in under a day, but only 35 percent actually did during a real downtime event. That's not a rounding error. It's a 25-point gap between the number that shows up in an SLA conversation and what happens when a server actually goes down.
Close that gap and you protect the contract. Leave it open and the first real outage becomes the moment a client learns your recovery time objective was a projection dressed up as a commitment. More than half of the organizations in that same survey said they plan to switch backup vendors within the year, and the ability to actually test recovery was one of the top reasons cited.
Why Are Attackers Targeting Backups Directly?
Because a clean, tested backup is the one thing that lets a client refuse to pay a ransom. Veeam's 2025 Ransomware Trends Report found that 89 percent of organizations hit by ransomware had their backup repositories directly targeted, and attackers modified or deleted an average of 34 percent of those repositories once inside. Despite that, only 32 percent of respondents were using immutable backup repositories at all.
That's the mechanism your testing has to account for. A backup that happens to survive an attack is luck, not a plan. Sophos's 2026 State of Ransomware report found backup-based recovery climbed to 66 percent of encrypted-data cases, up 12 points from the year before, which is the right direction, but it also means roughly a third of victims still couldn't restore from backup when it mattered.
Low immutability adoption is the part worth sitting with. If two-thirds of organizations still don't run immutable repositories, testing has to cover more than "did the backup complete." It has to confirm the backup is actually isolated from the credentials and management consoles an attacker would use to reach it, because a backup an attacker can delete is not meaningfully different from no backup at all.
What Changes When Testing Becomes Continuous
Annual DR drills and quarterly tabletop exercises still have a place, but they're a snapshot, not a schedule. The bigger shift in 2026 is automated, non-disruptive testing that spins up isolated recovery environments to validate recovery time and recovery point targets against a live environment, on a running basis, and produces evidence you can hand a client or an auditor without staging a fire drill every time.
This is where operations matter more than the tooling. Automated verification is worthless if nobody owns reviewing what it flags, escalating failures, and updating runbooks when a client's environment changes underneath it. That's a process problem before it's a product problem, and it's exactly the kind of overhead a growing MSP tends to underbuild for as client count climbs faster than technician headcount. Provisioning and workflow tools like Catalyst exist because that overhead is real, and it compounds with every new client you add to the stack.
Think about what happens operationally when a new client comes on board mid-quarter. Their environment needs to enter your testing rotation on day one, not whenever someone remembers to add it. Without a defined intake step, new clients quietly sit outside the cadence for months, which is exactly the population most likely to have an untested recovery plan when something actually breaks.
Building a Testing Cadence That Holds Up
A cadence only works if it's written down and assigned to a specific role, not implied by whoever happens to notice a failed job. The table below is a starting structure, not a finished policy.
| Test type | Minimum frequency | Who owns it |
|---|---|---|
| Automated backup verification | Daily | Monitoring tooling, reviewed weekly by a technician |
| Targeted file or VM recovery | Monthly | Technician on rotation |
| Application-level tabletop exercise | Quarterly | Service delivery lead |
| Full DR failover test | Annually, minimum | Operations lead, with client sign-off |
Once the cadence is defined, the harder part is keeping it running as staff turn over and new clients get added faster than your documentation can keep pace. That's a stack decision as much as a policy one. If you haven't mapped which tools in your current lineup actually enforce this cadence versus which ones just assume a human will remember, the stack builder is a fast way to find that blind spot before a client does.
What This Costs You If You Skip It
Backup and recovery is the most commonly offered managed service in the channel, delivered by 79 percent of MSPs according to Kaseya's 2026 State of the MSP Report, and it's also climbing as a stated client priority, ranking third among the services clients say they need most this year. Roughly half of MSPs reported year-over-year revenue growth in BCDR specifically. That combination cuts both ways.
The revenue opportunity is real, but so is the scrutiny that comes with it. Clients are getting more comfortable asking how you know recovery will work, not just whether you sell it. The cost of skipping testing doesn't show up on an invoice. It shows up as a failed restore during the one outage that actually matters, a lost contract after a client discovers testing was assumed rather than done, or hours of unbillable technician time improvising a recovery that should have been rehearsed months earlier.
None of that shows up in a pipeline report either, which is why it's easy to underweight until it happens once. A single failed recovery during a real incident tends to cost an MSP more in churn and reputation than years of disciplined testing would have cost in labor hours. Treat testing as a margin decision, not a compliance chore, and the staffing time for it stops looking optional.
Testing discipline is what turns a backup product into an actual recovery guarantee, and it's one of the clearer ways to differentiate a BCDR offering that otherwise looks the same as every other line item on a competitor's price sheet. If you're rethinking how backup and recovery fit into your broader catalog, our full product lineup is built around that kind of operational discipline. See the full stack to see how the pieces fit together.
Sources: Datto State of BCDR Report 2025 | Veeam 2025 Ransomware Trends Report | Sophos State of Ransomware 2026 | Kaseya 2026 State of the MSP Report.