Skip to content
RansomwareBackup
recovery microsoft 365google workspaceslackatlassian

RTO vs RPO: what each one actually costs you

RTO is how long you can be down; RPO is how much work you can lose. What each means, and the real numbers behind the four SaaS platforms.

RTO vs RPO is the difference between how long you can be down and how much work you can afford to lose. RTO — recovery time objective — is the clock running forward from the incident to the moment people are working again. RPO — recovery point objective — runs backward: the last moment your data was clean, and everything after it is gone. They are separate numbers, they are bought with different things, and on Microsoft 365, Google Workspace, Slack and Atlassian Cloud most organisations have never set either. The platform set them, by default, and the numbers only become visible during an incident.

The two definitions, precisely

Both terms come from formal contingency-planning literature rather than from vendors, which is worth knowing when a vendor uses them loosely. The NIST glossary, citing NIST SP 800-34 Rev. 1, defines them as:

Recovery time objective: The overall length of time an information system's components can be in the recovery phase before negatively impacting the organization's mission or mission/business processes.

Recovery point objective: The point in time to which data must be recovered after an outage.

Two things follow that are easy to miss. First, both are objectives — statements of what the business requires, not descriptions of what your tooling happens to do. The point of writing them down is to compare them against reality and find the gap. Second, RTO is defined in terms of impact on the business, not in terms of technology. An RTO is not "how fast can we restore"; it is "how long before this genuinely hurts".

The practical translation:

  • RPO answers: if we recover, how much work vanishes? An RPO of 24 hours means you accept losing up to a day.
  • RTO answers: how long is the business degraded or stopped? An RTO of four hours means you accept being down half a working day.

They are bought separately. A tighter RPO is bought with more frequent recovery points. A tighter RTO is bought with a faster, better-rehearsed restore path. Money spent on one does not improve the other, which is why the pair get confused and why the confusion is expensive.

RTO vs RPO: why the distinction decides what you buy

Consider two organisations hit by the same ransomware incident.

The first runs a busy Exchange Online tenant where deals are negotiated over email. Losing four hours of mail is a serious commercial problem: they need a tight RPO. Being read-only for a day while things are rebuilt is survivable.

The second runs a support desk in Slack and Jira. Losing an hour of messages is an inconvenience. Being unable to work for a day is an outage their customers will notice and their service-level commitments may penalise: they need a tight RTO.

Same incident, different purchases. Getting this backwards is the most common planning error in this area — and note that a backup product only ever sells you one of the two directly. Recovery points are a feature. Recovery time is mostly a property of your restore path, your scale and your practice.

What Microsoft 365 actually gives you

Microsoft is the only one of the four platforms that publishes both numbers, and it deserves credit for doing so — its documentation for the Microsoft 365 Backup add-on even labels one row of the feature table "Restore speeds (RTO)".

On the RPO side, Microsoft documents recovery points of 10 minutes for two weeks prior, then weekly snapshots from 2 to 52 weeks prior for OneDrive and SharePoint, and 10 minutes for the prior 52 weeks for Exchange Online. Read that carefully, because the shape matters more than the headline. Inside two weeks your OneDrive RPO is ten minutes. At three weeks it is a week — recover to a weekly snapshot and you lose up to seven days of everything changed since. Exchange Online keeps ten-minute granularity for the full year, which is a materially better guarantee than the file workloads get.

The recovery window itself is configurable on a backup policy at 3 months, 6 months, 1 year or 2 years, and existing policies default to 1 year.

On the RTO side, Microsoft publishes median restore expectations rather than a promise: a single OneDrive or SharePoint unit in around 30 minutes with an express restore point (with a note that single-unit restores range between roughly 10 and 120 minutes depending on size), a single Exchange mailbox in around two hours, 250 units in three to four hours, and at 1,000-plus units up to 250 protection units per hour. Mailbox item recovery is quoted at roughly 100 to 500 items per minute.

Do the arithmetic for your own tenant before you assume you are covered. Two thousand mailboxes at those rates is not an afternoon. And there are documented ceilings: up to 1,000,000 items per workload, 100,000 items per backup policy.

One more detail that belongs in an RTO conversation and rarely appears in one. A OneDrive or SharePoint restore is described as a rollback that overwrites all content and metadata created since that prior point in time, with per-file restore via versions listed as coming soon. So the restore is site-shaped, not file-shaped: recovering yesterday's ransomware event may also discard this morning's legitimate work. That is not a flaw so much as a design, but it changes what "recovered" means, and it belongs in the plan rather than in the incident. We set out the add-on's wider boundaries in what Microsoft 365 backup actually covers.

The other three platforms do not give you an RPO at all

This is the part that surprises people, and it is a direct consequence of how native recovery works.

Native SaaS recovery is not point-in-time. It is a bin with a deadline. You cannot say "return this Google Drive to its 09:00 state"; you can restore what was deleted, if you are inside the window, and if you can identify it. That means your RPO is not a duration at all — it is an event. You recover to the moment before the deletion, or you do not recover.

Against ransomware, which overwrites rather than deletes, that distinction is decisive: an encrypted file is a new version of a file, not a deleted one, so there is often nothing in the bin to restore.

  • Google Workspace gives administrators a limited window to restore a deleted user's data, applied as a date range rather than per item, and Vault exports rather than restores. The four available routes and what each misses are covered separately.
  • Slack and Atlassian Cloud hand you an export — a point-in-time file with no restore path back into the product. The RPO is the moment you ran the export. The RTO is however long it takes a human to reconstruct a workspace by hand, which is not a number anyone can commit to.

If your disaster-recovery document states an RPO and an RTO for these workloads today, it is stating an aspiration. The honest version of that document either records the real numbers or records that an independent copy is required to reach the stated ones — which is the point where a ransomware disaster recovery plan stops being paperwork.

The trade-off nobody puts on the slide

Isolation and recovery time pull against each other, and Microsoft makes the argument itself: it markets the speed of its in-boundary backup against the alternative of "copying data at a scale from a remote, air-gapped location requiring weeks or even months to get your business back up and running."

Treat the framing with appropriate scepticism — it is a vendor arguing for its own architecture — but the underlying physics is real. A copy held close to the tenant restores quickly and shares the tenant's fate. A copy held at a distance survives the tenant's fate and takes longer to bring home. That is precisely the tension in what an air-gapped backup actually cuts, viewed from the RTO side.

Which is why the two numbers have to be set together. An organisation that optimises only for RTO ends up with fast copies inside the blast radius. One that optimises only for isolation ends up with a pristine copy and a two-week restore. Neither is a plan.

Setting numbers that mean something

  1. Set them per workload, not per company. Mail, files, chat and issue trackers have genuinely different tolerances. One company-wide RTO is a number nobody owns.
  2. Set them from business impact, not from your tooling. Start with when the pain becomes real; then measure the gap to what you have. Deriving the objective from the current capability guarantees you always meet it and learn nothing.
  3. Write down today's actual numbers next to the targets. For most SaaS-native estates the honest entry in the RPO column is "last deletion event" and in the RTO column "unknown".
  4. Count your units, not your gigabytes. Published restore rates scale with the number of sites, accounts and mailboxes. Your worst case is a tenant-wide event, so size it that way.
  5. Test the RTO, because it is the only one you can get wrong quietly. An RPO can be verified from documentation. An RTO is a claim about your organisation under pressure, and the only evidence is a rehearsal.

Where we fit

We are not the platform and we do not run your tenant. What we do is turn this into two numbers per workload that your business actually agrees with — what your recovery point and recovery time have to be, what the native windows give you today, and where the gap is — then help you choose the right fit for your stack and get it deployed and running. You own and operate it from there.

If your current RTO and RPO are assumptions rather than measurements, a recovery assessment walks your workloads, the documented native windows behind each one, and what a realistic tenant-wide restore would take.