Cloud architecture

Managed Services Move Reliability Work

Cloud abstraction changes reliability responsibility rather than removing it.

Jason Doyle First published: 15 August 2025 27 minute read

Disclosure: These views are my own and do not represent my current or any former employers.

Executive summary

Managed cloud services are often described as a way to stop operating undifferentiated infrastructure. That description is partly true. A team that adopts a managed database, queue, identity system, load balancer, Kubernetes control plane, analytics platform, or software-as-a-service product no longer patches the same machines, replaces the same disks, or writes the same failover automation it would have needed in a self-operated system.

The mistake is to treat that as the end of reliability work.

Managed services move reliability work. They remove some component operation and replace it with dependency management, architecture, verification, contract review, support readiness, recovery planning, and exit judgement. The provider takes responsibility for parts of the stack. The customer still owns the reliability of the workload that depends on that stack.

The public cloud providers say this in their own language. AWS says customers remain responsible for their data, classifications, permissions, and service configuration even when AWS operates abstracted services such as S3 and DynamoDB.[1] Microsoft says cloud customers always retain responsibility for data, endpoints, accounts, and access management, and that configurations and settings remain the customer's responsibility across cloud models.[2] Google says customers must understand each service, each configuration profile, and the controls they need for their workloads, while Google argues for a broader model of "shared fate" rather than a narrow transfer of responsibility.[3]

The same logic applies to reliability. The provider may run the database engine. The customer must still decide whether the database is the right failure domain, whether its quota is sufficient, whether restore time meets the business objective, whether a region outage is tolerable, whether the control plane is needed during recovery, whether telemetry is independent enough, and whether the organisation can leave the service if the risk changes.

Public incidents make this concrete. AWS's 7 December 2021 US-EAST-1 event showed data-plane work continuing in some places while control planes, monitoring, support, and recovery workflows were impaired.[11] AWS's 2020 Kinesis event and Google's 2021 load-balancing incident show dependency fan-out and configuration-pipeline risk.[12][13]

These are not arguments against cloud providers. They are evidence that managed services have failure modes that are different from self-operation. Some failures affect control planes rather than data planes. Some affect global services rather than a single zone. Some affect monitoring, support, or recovery workflows at the same time as the service. Some arise from service updates, configuration systems, quota limits, or latent dependencies that customers cannot inspect directly.

The strongest counterargument is important: managed services often improve real reliability and security. Most organisations cannot operate physical datacentres, replicated storage, fleet patching, DDoS protection, key management, database backups, capacity planning, and 24-hour incident response at the standard of a major cloud provider. Azure's shared-responsibility documentation explicitly notes that on-premises environments often leave responsibilities unmet, including delayed patching, inadequate physical security, incomplete monitoring, outdated hardware, and insufficient backup and disaster recovery.[2] The right conclusion is not to reject managed services. It is to stop treating adoption as delegation without remainder.

This paper proposes a managed-service reliability review. Before relying on a service, the organisation should record what responsibility moved to the provider, what responsibility remains with the customer, which operations are control-plane dependent, which limits and quotas can fail the workload, how the service behaves across zones and regions, what backups and restores actually prove, what telemetry is independent, what the SLA excludes, what support path exists during an outage, and what exit would require.

Avoid simplistic multi-cloud advocacy. Running the same workload across providers can add cost, latency, security complexity, inconsistent semantics, weaker operational focus, and new failure modes. Concentrating on one provider is rational when the provider materially improves reliability, the service is not strategically constraining, the business can tolerate the failure domain, and the portability tax would reduce more important work. Portability is worth its cost when regulation, customer commitments, recovery objectives, bargaining power, or business survival require a credible alternative.

Managed services do not remove reliability responsibility. They change where that responsibility must be exercised.

1. The work did not disappear

A managed service changes the operator of a component. It does not change the fact that the component participates in a system.

That distinction sounds obvious until a system is being designed under time pressure. A team chooses a managed database because it does not want to operate a database. A team chooses a hosted identity provider because it does not want to build authentication. A team chooses a serverless queue because it does not want to run brokers. The decision is often sensible. It may remove years of patching, backup scripting, replication tuning, failover testing, and hardware planning.

But the application still depends on a database, identity provider, queue, or broker. The user does not experience a provider responsibility matrix. The user experiences whether the product works.

Reliability work therefore reappears in a different form:

  • selecting the right service and tier;
  • understanding the service's failure domains;
  • configuring redundancy that is not automatic;
  • requesting and monitoring quotas;
  • testing retry, timeout, and back-pressure behaviour;
  • separating control-plane assumptions from data-plane assumptions;
  • proving backup and restore under realistic conditions;
  • arranging support and escalation;
  • measuring user-visible service levels;
  • deciding how much lock-in is acceptable.

These tasks are less visible than replacing disks, tuning kernels, or maintaining database replicas. They are still operations.

The risk is a false binary. Either the team "runs" the component or the provider does. In practice, reliability sits at the boundary between them. The provider operates the service. The customer operates the dependency.

A managed queue can be highly reliable while the application exhausts account-level throughput. A managed database can preserve storage while the customer cannot restore quickly enough. A managed Kubernetes control plane can be patched while workloads still fail because maintenance windows, disruption budgets, or node repair behaviour were misunderstood. A secure identity provider can still leave the customer without an emergency access path.

The work changes from component mechanics to system judgement.

2. Shared responsibility is also a reliability model

Cloud shared-responsibility documents are usually written for security and compliance. Their structure is useful for reliability because they expose the same boundary problem.

AWS says responsibility is determined by the services a customer selects. In EC2, the customer manages the guest operating system, applications, and security groups. In abstracted services such as S3 and DynamoDB, AWS operates the infrastructure, operating system, and platform, while customers manage data, classifications, encryption options, and permissions.[1] Microsoft separates responsibility by IaaS, PaaS, and SaaS, but says customers always retain data, endpoints, accounts, and access management. It also states that configurations and settings are the customer's responsibility across deployment types.[2] Google says the shared-responsibility model can be hard because customers must understand each service, each configuration option, and what Google does to secure the service.[3]

Translate those statements from security to reliability:

Security responsibility Reliability analogue
Customer owns data classification Customer owns data criticality and loss tolerance
Customer configures access Customer configures redundancy, scaling, and failover
Provider secures physical hosts Provider operates the service infrastructure
Customer chooses encryption and keys Customer chooses recovery points, key availability, and restore access
Customer monitors compliance Customer monitors user-visible reliability and provider dependency health

The boundary is not the same for every service. A managed object store, managed database, hosted CI system, identity provider, serverless function platform, and SaaS collaboration suite all expose different controls and different failure modes.

A useful review therefore begins with a responsibility map, not a product name.

Ask what the provider does, what the provider exposes, what the customer can configure, what the customer can observe, and what neither party can guarantee alone. The answer may differ by region, tier, SKU, edition, support plan, or optional feature.

This is especially important for services that advertise high availability only when the customer chooses a particular configuration. Azure's availability-zone documentation says some services automatically use multiple zones, while others require the customer to configure multi-zone deployment. Zonal resources do not automatically provide resilience to a zone outage; the customer must design resources in multiple zones and handle failover.[6] Similar distinctions exist across providers.

The practical rule is simple: if reliability depends on a configuration, it is still reliability work.

3. Control planes fail differently from data planes

One of the most useful reliability distinctions is between control plane and data plane.

AWS defines control planes as administrative APIs used to create, read, update, delete, and list resources. Launching an EC2 instance, creating an S3 bucket, and describing an SQS queue are control-plane actions. Data planes provide the primary function of a service: a running EC2 instance, reading and writing an EBS volume, getting and putting objects in S3, or Route 53 answering DNS queries.[4]

The distinction matters because recovery plans often depend on control planes without saying so.

A team may believe it has a regional failover plan because it can create replacement capacity in another region. That plan depends on instance-launch APIs, load balancer creation, DNS changes, certificate operations, identity permissions, image access, infrastructure-as-code state, and quota. If those actions rely on an impaired control plane, the plan may fail at the moment it is needed.

AWS's own fault-isolation guidance warns against relying on global service control planes in recovery paths. It notes that some global services have control planes in a single region while data planes are globally distributed. For example, IAM's control plane in the commercial AWS partition is in US-EAST-1, while its data plane is isolated and distributed regionally. Route 53 public DNS has a control plane in US-EAST-1 and a globally distributed data plane. AWS recommends relying on data-plane operations, not control-plane operations, during recovery.[5]

The 7 December 2021 AWS event showed why that advice matters. Existing EC2 instances were not directly affected in the same way as EC2 APIs used to launch or describe instances. Existing Route 53 DNS answers continued, while Route 53 APIs for changing DNS records were impaired. Existing load balancers remained healthy, while API errors delayed provisioning and instance registration. Support case creation was also affected for several hours.[11]

This is not evidence that one provider is uniquely vulnerable. Google Cloud's 2021 GCLB incident showed a different version of the same boundary. A configuration-pipeline bug caused global HTTP(S) load balancing endpoints to return 404 errors. During partial recovery, traffic was restored, but customer configuration changes were suspended while Google validated the fix and resumed pushes.[13]

For managed services, a reliability review should classify operations into at least four groups:

  1. Normal data-plane operations that must continue during an incident.
  2. Control-plane operations needed for deployment, scaling, or repair.
  3. Break-glass operations needed for recovery.
  4. Administrative operations that can wait.

The third group is the dangerous one. If a recovery plan says "create", "update", "attach", "promote", "rotate", "increase", "register", or "change DNS", it probably depends on a control plane. The safer design is to pre-provision what must exist in a disaster, test that it works without last-minute creation, and make failover depend on the smallest possible set of highly available data-plane actions.

4. Quotas and limits are reliability constraints

Quotas are sometimes treated as billing or governance details. They are also reliability constraints.

AWS says accounts have default quotas for each service, that many quotas are region-specific, that some can be increased, and that not all can be increased. It also says support may approve, deny, or partially approve increase requests, and that increases are not immediate.[8] Azure documents subscription and service limits, including adjustable and non-adjustable limits, and notes that some limits are managed at a regional level. A vCPU quota increase for West Europe does not increase quota in other regions.[9] Google Cloud says quotas restrict resources such as API calls, load balancers, projects, and hardware allocation; when a task exceeds quota, access is blocked and the task fails. Google distinguishes allocation, rate, and concurrent quotas, and global, regional, zonal, project, folder, organisation, and user-level quotas.[10]

A quota failure can look like an outage to the application:

  • failover cannot create enough instances in the recovery region;
  • autoscaling cannot add capacity during a traffic spike;
  • a managed queue or API returns throttling errors;
  • a deployment pipeline is rate-limited while trying to repair an incident;
  • a backup, restore, export, or support operation exceeds a hidden or poorly understood limit.

Quotas are especially dangerous in disaster recovery because normal usage does not prove recovery capacity. A warm standby may need to absorb traffic from a failed region. A pilot-light environment may need to scale from a small footprint to production load. A backup-and-restore plan may need to create large databases, network paths, load balancers, and caches in a region where the account normally uses little capacity.

The managed-service review should therefore record:

  • default quota;
  • current approved quota;
  • observed peak usage;
  • failover usage;
  • whether the quota is global, regional, zonal, or per project;
  • whether the quota is adjustable;
  • lead time for increase;
  • owner of quota monitoring;
  • alert threshold;
  • tested behaviour at quota exhaustion.

Do not wait for a disaster to discover that the recovery plan assumes a support case.

5. Regions, zones, and global dependencies are not interchangeable

Cloud reliability language can become vague. "Multi-AZ", "zone redundant", "regional", "multi-region", and "global" sound like layers of safety. They are only useful when they map to the workload's failure modes.

AWS says each region consists of multiple independent and physically separate Availability Zones, and that regions are isolated and independent from other regions with exceptions for some global operations.[5] It describes Availability Zones as discrete datacentres with separate power, networking, and connectivity, meaningfully distant from each other but close enough for low-latency synchronous replication.[5] Azure describes availability zones as separated groups of datacentres with independent power, cooling, and networking, but also says different services support zones differently. Some resources are zone redundant. Some are zonal. Some are nonzonal or regional.[6]

A managed service may hide some of this detail. That can be a benefit. It can also hide assumptions.

A service may be "regional" but depend on a global identity, DNS, certificate, quota, telemetry, billing, deployment, or support component. A service may be zone redundant only in certain regions or tiers. A database may support cross-region replicas but not synchronous writes across arbitrary distances. A global load balancer may improve user routing while creating a singleton configuration dependency.

Google's multi-regional deployment guidance is careful about this trade-off. It recommends multi-regional deployment for business-critical applications where region-outage tolerance is essential, but notes that cross-location traffic, replicated data, and operational complexity increase cost and complexity.[7] AWS disaster-recovery guidance similarly frames backup and restore, pilot light, warm standby, and active-active multi-region as choices with increasing cost and complexity and decreasing RTO and RPO.[22]

The more useful question is "What is the smallest failure domain that can break the user journey?" rather than "Is this service managed?"

For each critical flow, identify:

  • the user-visible operation;
  • every managed service on the path;
  • whether each service is zonal, regional, multi-region, or global;
  • whether the configured tier changes that answer;
  • whether identity, DNS, keys, secrets, logging, metrics, and deployment have the same failure domain;
  • whether recovery requires new control-plane operations;
  • whether data consistency requirements prevent fast failover.

A single-region deployment can be a rational choice for a workload with modest availability needs, strong data consistency needs, low recovery urgency, or a small team. A multi-region design can be necessary for a critical service. It can also be a way to spend reliability effort on the wrong problem while leaving identity, observability, backups, or customer communications as single points of failure.

6. Updates, opacity, and support are part of the dependency

Managed services improve reliability partly because providers update them. That also means the customer depends on provider change management.

Amazon RDS maintenance can include underlying hardware, operating system updates, and database engine versions. Some maintenance items require the DB instance to be offline briefly. RDS says required patching is automatically scheduled for security and reliability patches, and that maintenance operations are not guaranteed to finish before the maintenance window ends. Multi-AZ deployments can reduce some operating-system maintenance impact, but engine upgrades may still make primary and secondary unavailable during the upgrade depending on engine and configuration.[14]

Cloud SQL automatically updates instances for hardware, operating system, and database engine reliability, performance, and security. Some updates require a brief interruption. Maintenance settings can control timing, but Cloud SQL's own documentation lists prerequisites, constraints, notification limits, and cases where downtime may be higher.[15] GKE maintenance policies provide windows and exclusions for some automatic maintenance, but Google says they do not block all maintenance. Control-plane repairs, critical security patching, and maintenance of dependent services may ignore windows because not doing them could leave clusters non-functional or vulnerable.[16]

These are reasonable provider choices. Security patches and repair operations cannot always wait for a customer's perfect window. The reliability point is that managed-service updates belong in the application reliability model.

The customer usually cannot inspect the implementation deeply. Public postmortems often reveal internal coupling only after an incident. AWS's 2020 Kinesis report described front-end shard-map construction, operating system thread limits, a latent Cognito buffering bug, CloudWatch's reliance on Kinesis, Lambda metric buffering, and EventBridge backlog processing.[12] Google described a GCLB configuration-pipeline race condition that existed for months before it manifested during a rollout.[13] These details are useful precisely because they were not details customers could have fully known beforehand.

Opacity need not imply distrust, but it does mean the customer needs controls that do not require perfect knowledge:

  • design for documented failure domains, not assumed internals;
  • avoid recovery paths that require creating new resources during an outage;
  • read service health and postmortems, but test locally;
  • subscribe to maintenance notifications and release notes;
  • stage provider upgrades where the service allows it;
  • test client behaviour during short connection loss;
  • keep independent evidence when provider telemetry may lag;
  • establish support paths before the incident.

Support is itself a dependency. During the 2021 AWS event, AWS said the Support Contact Center relied on the internal AWS network and that creating support cases was impacted from 7:33 AM until 2:25 PM PST.[11] A support plan is not a recovery plan if the recovery plan assumes support will be reachable, informed, and able to change the situation within the workload's RTO.

7. Backup and restore are managed, but recovery is owned

Managed services often include backup features. That is not the same as a tested recovery capability.

Amazon RDS creates automated backups and can recover a DB instance to a point in time within the retention period. But after restore, volumes can continue loading data blocks from S3 in the background, so the instance may be available before performance is fully initialised. Some engines have special restore considerations, and some actions can break point-in-time recovery sequences.[17]

Azure SQL Database creates full, differential, and transaction-log backups and supports point-in-time restore, long-term retention, and geo-restore depending on configuration. Storage redundancy choices matter: locally redundant storage is not recommended for regional-outage resilience, and changes apply only to future backups.[18]

Cloud SQL supports on-demand, automated, retained, and final backups. Backups can support disaster recovery by creating a new instance in another region or zone, but replicas do not have their own backups until promoted, and backup and restore cannot upgrade a database to a later version.[19]

These features are valuable. They still require customer decisions:

  • Is point-in-time recovery enabled?
  • What is the actual latest restorable time?
  • Are backups protected from accidental deletion and malicious use?
  • Are keys, identity, network paths, and target quotas available during restore?
  • Does restore create a new resource that must be reattached to applications?
  • Does performance after restore meet the business objective?
  • Are replicas, exports, and backups being confused?
  • Has a full restore been tested at production scale?

Atlassian's April 2022 cloud outage is a useful public case study. Atlassian said approximately 400 cloud customer sites were improperly deleted after a script was run with the wrong execution mode and wrong IDs. The company maintained immutable backups and routinely restored individual customers or small groups. What it had not automated was restoring a large subset of customers into the existing, in-use environment without affecting other customers. Recovery required manually extracting and restoring pieces from backups, validating each site, and working with customers. Some batches took four to five elapsed days to hand back.[28]

Backups did not fail in that case. The material issue was restore shape: a backup that can restore one tenant, all tenants, or a new environment may not support restoring a large subset into a live shared environment quickly.

GitHub's 2018 incident provides another recovery lesson. GitHub had MySQL backups every four hours, retained for years, and tested daily. During the incident, it still took hours to restore multiple terabytes, transfer data from remote blob storage, decompress, checksum, prepare, and load backups. GitHub chose data integrity over faster restoration of usability.[29]

A reliability review should therefore treat "backup exists" as the beginning of the question. The evidence is a timed, audited, end-to-end restore that proves the RTO, RPO, data integrity, access path, and application reconnection procedure.

8. Observability boundaries and game days

A managed service can expose metrics, logs, traces, audit events, health feeds, and support notifications. It rarely exposes everything the provider sees.

That boundary matters during incidents. AWS's December 2021 event impaired internal monitoring, delayed AWS's own understanding, affected CloudWatch monitoring for customers, and left some metrics missing for parts of the event.[11] The November 2020 Kinesis event affected CloudWatch metrics and alarms, causing alarms to move to INSUFFICIENT_DATA and creating gaps in CloudWatch metrics.[12]

A team that relies only on provider-native telemetry may lose visibility when the provider service, its telemetry pipeline, or a shared dependency is degraded. The answer is not to duplicate every provider system. It is to make sure the most important user-visible signals are independent enough to guide response.

For critical managed-service dependencies, monitor at several layers:

  • synthetic checks from outside the provider account or region;
  • client-side success rate, latency, and timeout behaviour;
  • provider metrics and health events;
  • quota usage and throttling;
  • backup age and restore test results;
  • control-plane API error rates where available;
  • business flow completion as well as component health.

SRE literature is useful here because it begins with user-visible objectives. Google's SRE book distinguishes service level indicators, objectives, and agreements, and argues that SLIs should directly measure the service level of interest where possible.[20] It also frames reliability as an explicit risk trade-off, not a pursuit of 100 per cent availability at any cost.[21]

Managed services do not change that discipline. They make it more important. The provider's SLA is not necessarily the customer's SLO. A workload with five managed dependencies can miss its user SLO even if no single provider SLA is breached.

Testing must include dependency failure. AWS Well-Architected recommends chaos experiments and resilience testing, including experiments for dependency latency, throttling, packet loss, DNS failures, and failover mechanisms. It explicitly warns against designing for resilience without verifying how the workload functions as a whole when faults occur.[23] NIST contingency-planning guidance identifies contingency planning, incident response, disaster recovery, testing, training, and exercises as part of resilience practice.[24]

Game days for managed services should not be theatrical. They should answer concrete questions:

  • What if the data plane works but the control plane is unavailable?
  • What if the service throttles writes for thirty minutes?
  • What if quota increase is denied during a launch?
  • What if provider metrics are delayed?
  • What if backups restore but the restored resource has a new endpoint?
  • What if a region is down and identity changes cannot be made?
  • What if support cannot be reached for two hours?
  • What if the provider applies emergency maintenance outside the preferred window?

The result should be a changed system, not a slide deck.

9. SLAs and contracts are not availability

Service-level agreements are useful. They are also narrow.

Google's SRE book distinguishes an SLO from an SLA by asking what happens if the objective is not met. If there is no explicit consequence, it is probably an SLO rather than an SLA.[20] In public cloud contracts, the consequence is often a service credit, not compensation for the customer's business loss.

Google's Cloud SQL SLA states that if Google does not meet the service level objective and the customer meets its obligations, the customer is eligible for financial credits. It also states that the SLA is the customer's sole and exclusive remedy for failure to meet the SLO.[26] Google's Compute Engine SLA says the customer must request financial credit within 60 days and provide log files showing downtime, that maximum aggregate credits are capped by the amount due for the covered services in affected regions, and that exclusions include factors outside Google's reasonable control, customer software or hardware, violations of the agreement, and quotas applied by the system or listed in the admin console.[26]

The details vary by provider and service. The pattern is common: the SLA does not make the customer whole, does not guarantee recovery within the customer's RTO, does not cover every dependency, and may require the customer to prove downtime.

A managed-service reliability contract inside the customer organisation should therefore be stricter than the provider SLA. It should define:

  • the user-visible SLO;
  • the provider SLA and its exclusions;
  • the internal error budget;
  • the support severity and response expectation;
  • evidence required to claim credits;
  • escalation contacts;
  • fallback operations;
  • who can accept residual risk.

Procurement should not own this alone. Engineering, security, legal, product, finance, and operations all see different parts of the risk. A low service credit may be acceptable for an internal reporting tool and unacceptable for a customer authentication path. The judgement depends on the workload, not the provider logo.

10. Portability, concentration risk, and exit planning

Lock-in is not automatically bad. Lack of exit judgement is bad.

Managed services create value by exposing higher-level capabilities. The same abstraction that saves time can make exit harder. A database's replication model, a queue's ordering semantics, an identity provider's policy language, a serverless platform's event model, a data warehouse's SQL dialect, and a SaaS product's workflow rules become part of the application.

NIST's cloud synopsis and recommendations identify cloud computing concerns including portability, interoperability, and security. Those concerns remain practical, not theoretical.[25] The UK Competition and Markets Authority's 2025 cloud services final decision found that market concentration, barriers to entry and expansion, and barriers to switching and multi-cloud affected competition in UK cloud services. It described technical and commercial barriers, including egress fees, latency between clouds, lack of transferable skills, and insufficient transparency on mitigating technical barriers.[27]

Those findings do not prove that every workload should be multi-cloud. They show that exit and switching are real economic and technical issues.

Concentration is rational when:

  • one provider's managed services materially reduce operational risk;
  • the team can understand and test the chosen failure domains;
  • the workload's availability need fits the provider and region design;
  • portability work would displace more valuable reliability work;
  • specialised provider features create product value;
  • data gravity makes cross-provider operation slower or riskier;
  • the organisation has enough bargaining power or contractual protection.

Portability is worth paying for when:

  • a single provider outage would threaten the business;
  • regulation or customer contracts require an exit path;
  • the service is strategically central and hard to replace;
  • pricing, support, or roadmap risk is material;
  • the provider's failure domain does not match the workload's obligations;
  • the organisation expects merger, divestiture, sovereign, or procurement constraints;
  • historical restore and export tests show that exit would otherwise be too slow.

There are middle positions between naive lock-in and full active-active multi-cloud:

  • use managed services but define data export formats;
  • keep infrastructure as code and documented rebuild procedures;
  • maintain identity federation that can survive provider change;
  • avoid unnecessary proprietary features in low-value areas;
  • isolate provider-specific adapters behind clear interfaces;
  • test periodic restore into a neutral environment;
  • negotiate assistance, data return, deletion, and transition clauses;
  • keep a current inventory of managed-service dependencies.

Exit planning is a way to know the price of the dependency, not a declaration of distrust.

11. The strongest counterargument

The strongest counterargument is that managed services usually make systems better.

For many organisations, self-operation is not a noble expression of control. It is a backlog of unpatched hosts, ageing hardware, ad hoc backups, tribal knowledge, weak monitoring, incomplete disaster recovery, and a small team expected to match the resilience of a global provider.

Azure's shared-responsibility documentation makes that case directly. It says cloud can help solve longstanding security challenges because on-premises organisations often have unmet responsibilities and limited resources. It lists delayed patching, inadequate physical security, incomplete network monitoring, outdated hardware, and insufficient backup and disaster recovery as common examples.[2]

The same is true for reliability. Major providers can invest in physical security, redundant power, global networks, capacity planning, specialised operations, DDoS defence, automated repair, fleet-wide patching, and service teams beyond the budget and scale of a single customer. Provider-managed databases and storage systems often have better durability, patching, monitoring, and routine failover than the self-managed systems they replace.

The counterargument is strongest for common components where the provider's operational excellence is clearly superior and the customer's differentiation is low. Few product teams should run their own object storage. Few should build a bespoke identity stack. Few should maintain a database platform from bare metal if a managed database meets their needs. Time not spent on commodity operation can be spent on product reliability, customer experience, security review, and incident practice.

There is also a security-reliability connection. A service that is patched promptly is more reliable against security-driven outages. A backup system with immutability can improve recovery from ransomware or accidental deletion. A provider's standardised controls can reduce the variance that causes incidents.

This paper accepts that argument.

Better component operation is not the same as completed system responsibility, and the two should not be confused.

A managed service can be the right answer and still require a dependency review, a recovery test, quota planning, independent monitoring, and an exit record. In fact, the more important the managed service is, the more important those practices become.

12. What this paper does not claim

This paper does not claim that managed services are unreliable. Many are more reliable than realistic self-operated alternatives.

It does not claim that self-operation gives real control. A team can control a component badly, slowly, or without enough staff.

It does not claim that every workload needs multi-region or multi-cloud architecture. Those designs can add cost, latency, inconsistency, security complexity, and operational risk.

It does not claim that public postmortems prove a provider is generally unsafe. They are included because they show real failure modes, recovery constraints, and dependency shapes.

It does not claim that provider SLAs are useless. They are useful contractual signals, but they are not the same as the customer's user-visible reliability objective.

It does not claim that all lock-in is bad. Some concentration is a rational purchase of capability. The claim is that the concentration should be deliberate, priced, reviewed, and reversible where the business requires reversibility.

The claim is this:

Managed services remove some component operation. They do not remove reliability responsibility. They exchange direct control for dependency risk, quotas, control-plane and regional failure modes, opaque implementation, support escalation, recovery constraints, and exit complexity. Responsible adoption means managing that exchange explicitly.

Conclusion

Cloud managed services are one of the most important reliability improvements available to modern engineering teams. They let small teams use infrastructure, databases, messaging, identity, analytics, and security capabilities that would once have required large specialist organisations.

That benefit is real.

So is the trade.

When a team adopts a managed service, it gives up some direct control and gains a dependency. The provider operates more of the stack. The customer must understand how the dependency behaves, how it fails, how it is configured, how it is observed, how it is restored, how it is supported, and how it could be replaced.

Reliability therefore moves upward. It moves from disks, hosts, and patch scripts to contracts, quotas, failure domains, control planes, recovery evidence, telemetry boundaries, and business risk. This is not lesser work. It is the work that determines whether a managed component becomes a reliable system.

The useful question is not whether the service is managed.

It is whether the organisation has managed the responsibility that remains.

About the author

Jason Doyle writes about reliable software, observability, applied AI, and practical controls for systems that influence human decisions. He publishes at jasondoyle.ie and can be contacted at contact@jasondoyle.ie.

References

  1. AWS, Shared Responsibility Model, https://aws.amazon.com/compliance/shared-responsibility-model/.
  2. Microsoft Learn, Shared responsibility in the cloud, archived source revision, 30 September 2024, https://github.com/MicrosoftDocs/azure-docs/blob/74af00eadd7acf31cf8ab22c90829c39d52fe100/articles/security/fundamentals/shared-responsibility.md.
  3. Google Cloud, Shared responsibility and shared fate, https://cloud.google.com/architecture/framework/security/shared-responsibility-shared-fate.
  4. AWS, Control planes and data planes, https://docs.aws.amazon.com/whitepapers/latest/aws-fault-isolation-boundaries/control-planes-and-data-planes.html.
  5. AWS, AWS Fault Isolation Boundaries, https://docs.aws.amazon.com/whitepapers/latest/aws-fault-isolation-boundaries/.
  6. Microsoft Learn, Azure Availability Zones and Mission-critical architecture pattern, https://learn.microsoft.com/en-us/azure/reliability/availability-zones-overview and https://learn.microsoft.com/en-us/azure/well-architected/mission-critical/mission-critical-architecture-pattern.
  7. Google Cloud, Multi-regional deployment archetype, https://cloud.google.com/architecture/deployment-archetypes/multiregional.
  8. AWS, AWS service quotas, https://docs.aws.amazon.com/general/latest/gr/aws_service_limits.html.
  9. Microsoft Learn, Azure subscription and service limits, quotas, and constraints, https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/azure-subscription-service-limits.
  10. Google Cloud, Quotas overview, https://cloud.google.com/docs/quotas/overview.
  11. AWS, December 7, 2021 AWS Service Event in US-EAST-1, https://aws.amazon.com/message/12721/.
  12. AWS, Amazon Kinesis Service Event in US-EAST-1, 25 November 2020, https://aws.amazon.com/message/11201/.
  13. Google Cloud Status, Google External Proxy Load Balancing incident, 16 November 2021, https://status.cloud.google.com/incidents/6PM5mNd43NbMqjCZ5REh.
  14. AWS, Maintaining a DB instance, https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_UpgradeDBInstance.Maintenance.html.
  15. Google Cloud, About maintenance on Cloud SQL instances, https://cloud.google.com/sql/docs/mysql/maintenance.
  16. Google Cloud, GKE maintenance windows and exclusions, https://cloud.google.com/kubernetes-engine/docs/concepts/maintenance-windows-and-exclusions.
  17. AWS, RDS backups and point-in-time restore, https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_WorkingWithAutomatedBackups.html and https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_PIT.html.
  18. Microsoft Learn, Azure SQL Database automated backups, https://learn.microsoft.com/en-us/azure/azure-sql/database/automated-backups-overview.
  19. Google Cloud, Cloud SQL backups, https://cloud.google.com/sql/docs/mysql/backup-recovery/backups.
  20. Google, SRE: Service Level Objectives, https://sre.google/sre-book/service-level-objectives/.
  21. Google, SRE: Embracing Risk, https://sre.google/sre-book/embracing-risk/.
  22. AWS Well-Architected, Use defined recovery strategies, https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_planning_for_recovery_disaster_recovery.html.
  23. AWS Well-Architected, Test resiliency using chaos engineering, https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_testing_resiliency_failure_injection_resiliency.html.
  24. NIST, Contingency Planning Guide for Federal Information Systems, SP 800-34 Rev. 1, https://csrc.nist.gov/pubs/sp/800/34/r1/final.
  25. NIST, Cloud Computing Synopsis and Recommendations, SP 800-146, https://csrc.nist.gov/pubs/sp/800/146/final.
  26. Google Cloud, Cloud SQL SLA and Compute Engine SLA, https://cloud.google.com/sql/sla and https://cloud.google.com/compute/sla.
  27. UK Competition and Markets Authority, Cloud services market investigation: final decision, 31 July 2025, https://www.gov.uk/cma-cases/cloud-services-market-investigation.
  28. Atlassian, April 2022 outage update, https://www.atlassian.com/blog/how-we-build/april-2022-outage-update.
  29. GitHub, October 21 post-incident analysis, https://github.blog/news-insights/company-news/oct21-post-incident-analysis/.