> Disclosure: These views are my own and do not represent my current or any former employers.

# Reliability Customers Can Use

## Turning service health into decision-grade product information

Author: Jason Doyle

First published: 15 March 2026

Source review updated: 25 August 2026

## Executive summary

Reliability work has become much better at describing systems. Teams define SLIs and SLOs, defend error budgets, page on burn rates, test recovery, and write post-incident reviews. Those practices matter because they turn reliability into engineering decisions about when to ship, pause, repair, or redesign.

Customers need a different surface. They do not operate the provider's internal service. They operate their own business processes on top of it. During an incident they need to decide whether to retry, fail over, pause work, switch channels, notify their own users, reconcile later, or wait. A green or red component rarely gives enough information for that decision.

The claim of this paper is that reliability becomes a product surface when customers can use it to make better decisions. Internal health metrics describe the provider's system. Customer reliability information describes the customer's usable options.

Google SRE defines SLIs and SLOs as carefully specified measurements and targets, while warning that user-relevant measurements are often harder to obtain than server-side proxies \[1\]. The SRE Workbook treats SLOs and error budgets as decision tools rather than decorative dashboards \[2\]. That discipline should extend beyond the operations room.

Public cloud providers already show the direction. AWS Health provides service and resource events, account-specific visibility, alerts, and API access for supported plans \[8\]\[9\]. Google Cloud positions Personalized Service Health as the primary channel for project-relevant incidents, with the public dashboard as a broad fallback \[10\]\[11\]. Azure separates global status, personalised Service Health, and Resource Health for individual resources \[12\]\[13\].

The public record also shows the difficulty. AWS's December 2021 US-EAST-1 event impaired internal monitoring, CloudWatch visibility, support case creation, and status tooling while customers were trying to understand exposure \[18\]. GitHub's October 2018 incident prioritised data integrity over faster recovery \[19\]. Fastly's June 2021 outage was detected within one minute and mostly recovered within 49 minutes, but still affected customers globally \[20\]. Cloudflare and Slack have both published incidents where dependencies and response tooling shaped recovery \[21\]\[22\].

This paper proposes a customer reliability record: a structured, versioned account of service health that gives scope, confidence, customer impact, recommended action, data and security statements, recovery evidence, and next update time. It does not expose raw telemetry, promise perfect knowledge, replace customer monitoring, or turn every outage into a legal report.

Raw transparency can be harmful. Logs may contain tokens, personal data, internal network names, file paths, and commercially sensitive information; OWASP warns that such material should not be logged or exposed directly \[30\]. Decision-grade reliability is therefore not maximum disclosure. It is selected disclosure with scope, confidence, accessibility, security review, and an operational purpose.

## 1. The gap between health and usable reliability

A system can be internally healthy and still fail a customer. It can also be internally impaired while a particular customer can continue safely.

Internal dashboards are built around ownership. They ask whether a cluster is saturated, whether a queue is growing, whether a deployment is failing, whether a control plane can create resources, or whether a dependency is returning errors. Those are necessary questions for responders.

A customer asks different questions:

| Customer question | Decision behind it |
| --- | --- |
| Can users still complete checkout? | Keep accepting orders or switch to a degraded flow |
| Are writes durable? | Continue processing or pause mutations |
| Are events delayed or lost? | Reconcile later or replay from source |
| Is this only one region? | Fail over or keep traffic local |
| Is authentication affected? | Notify employees to stop retrying login |
| Is recovery stable? | Reopen the workflow or wait longer |
| Is my tenant affected? | Communicate to my own customers |

A traditional status page often compresses all of this into one label. Operational, degraded performance, partial outage, major outage, or maintenance can be accurate at platform level and still insufficient for customer action.

The problem is not that status pages are useless. They are valuable because they establish a public source of truth, provide subscription channels, and reduce rumours. Atlassian's incident communication guidance tells teams to acknowledge issues quickly, summarise known impact, update at a defined cadence, use plain language, stay consistent across channels, and own the problem even when a third-party provider is involved \[14\]. That advice is sound.

The gap is that many enterprise customers need a reliability answer at the level of their own dependency graph. A bank, hospital, manufacturer, public agency, or software company may depend on the same SaaS or cloud service in very different ways. One customer may be unable to sign new users in. Another may be able to serve existing sessions but unable to provision new infrastructure. A third may be unaffected because it uses a different region or feature set.

A component-level page cannot carry all of that meaning. A customer-level record can.

## 2. Internal SLOs are necessary but not sufficient

SLO practice begins in the right place: the user. The SRE literature advises teams to choose indicators that approximate what users actually experience, to avoid false precision, and to treat 100 per cent reliability as the wrong target for most services \[1\]\[2\]. It also explains that an SLO without an error budget policy becomes another metric rather than a decision-making tool \[2\].

That is the first important lesson for customer reliability. A number is useful only if it changes a decision.

Internal SLOs often decide engineering behaviour:

- page an on-call engineer;
- slow or stop releases;
- redirect work from features to resilience;
- declare an incident;
- prioritise a dependency fix;
- hold a postmortem when budget consumption is large.

Google's example error budget policy makes this explicit. If the service exceeds its error budget, releases stop except for severe issues or security fixes; if a single incident consumes more than a threshold, a postmortem is required \[3\]. The policy is a control surface for engineers and product leaders.

Customers need the same type of control surface, but the control is different. They usually cannot patch the provider's service. They can change their own behaviour. They can pause a batch, change routing, increase backoff, suppress noisy retries, prepare a customer notice, open a support case with the right evidence, or delay a migration.

The provider's SLO may not tell them which of those actions is rational.

### Aggregation hides local pain

A monthly availability target can be met while a small cohort has a bad day. A regional SLO can be green while a specific tenant, shard, identity provider, API method, marketplace integration, or data pipeline is impaired. A weighted global availability number can hide low-volume but high-consequence flows.

The SRE Workbook recognises this measurement problem when it distinguishes between an SLI specification, the user outcome that matters, and an SLI implementation, the available way of measuring it \[2\]. Server logs may miss requests that never reach the backend. Probers may miss a tenant-specific authorisation path. Client-side telemetry may be closer to experience but harder to collect consistently.

A customer reliability surface should therefore avoid saying more than the evidence supports. It should distinguish:

- observed impact to this customer;
- probable impact because the customer uses affected services;
- possible impact because a dependency is implicated but not confirmed;
- no known impact, with stated limits of detection;
- unknown, because the relevant signal is absent or delayed.

The word "unknown" is not a failure of communication. It is often the most honest state.

### SLA compliance is not operational guidance

An SLA answers a contractual question after a measurement period. An incident update answers an operational question now. The two may use related data, but they should not be confused.

Azure's reliability guidance makes the distinction clearly: an SLO is a measurable target based on the quality of service customers expect, while an SLA is a contractual agreement that may have financial or legal consequences \[6\]. It also emphasises that reliability targets should include availability, correctness, and recovery, and should be derived with business stakeholders \[6\].

A customer experiencing an outage cannot wait for the monthly SLA calculation to decide whether to fail over this morning. Decision-grade reliability has to operate at incident time; the credit-request calculation happens afterwards, on a different clock entirely.
## 3. Public evidence from service-health practice

The strongest evidence for customer-usable reliability comes from providers that already expose personalised health.

AWS Health gives customers event information about service and resource changes that may affect applications, with alerts, EventBridge integration, and account-specific events \[8\]. The AWS Health API lets eligible customers integrate health events into internal systems, use organisational views, and distinguish public from account-specific events \[9\].

Google Cloud makes a similar separation. Its public Cloud Service Health dashboard is an at-a-glance page for broad severe incidents. Personalized Service Health is the recommended primary channel for disruptions relevant to a customer's projects, available through a dashboard, API, Cloud Logging, and alerts \[10\]\[11\]. Google also recommends fallback strategies for cases where the personalised channel itself depends on affected services, such as identity \[11\].

Azure divides the surface into global status, Service Health for personalised service-impacting communications, and Resource Health for individual resource state. Resource Health can show whether a particular resource is Available, Unavailable, Unknown, or Degraded and may report actions Microsoft is taking as well as actions the customer can take \[12\]\[13\].

These practices point to a useful pattern:

| Surface | Primary audience | Main value | Limitation |
| --- | --- | --- | --- |
| Public status page | Anyone | Broad shared awareness | Too coarse for many customer decisions |
| Authenticated health dashboard | Affected customer teams | Relevance to subscriptions, projects, or resources | Depends on identity and provider mapping |
| Health API or event feed | Automation and SRE teams | Integration into runbooks, paging, and dashboards | Requires schema quality and stable semantics |
| Resource or tenant health | Service owners and support | Specific diagnosis and evidence | Harder to implement and explain |
| Post-incident report | Operators, executives, customers | Learning, accountability, corrective actions | Arrives after immediate decisions |

The pattern falls short of universal maturity, but it is evidence that customers already need more than a public red light.

## 4. What customers are trying to decide

A reliability update should be designed backwards from customer action.

| Decision | What the customer needs to know |
| --- | --- |
| Retry | Are requests safe to retry, should they use backoff, and is idempotency required? |
| Fail over | Is the impact regional, zonal, product-specific, or control-plane only? |
| Pause work | Are writes, migrations, exports, or scheduled jobs at risk? |
| Degrade gracefully | Which features can continue and which should be disabled? |
| Communicate impact | Which users, locations, records, or time windows are affected? |
| Wait | Is recovery underway, what is the next update time, and what evidence would change the advice? |

A useful health record does not expose the provider's entire incident bridge. It helps the customer choose among these options.

Retry guidance matters because customer behaviour can help or harm recovery. Aggressive retries may deepen an overload. Blind retries after ambiguous writes may create duplicates. The record should therefore distinguish safe retry, retry with backoff, do not retry, and verify before retrying.

Failover guidance must be even more careful. Failover can create data divergence, operational load, and customer confusion. A provider should not command every customer to fail over. It should give enough scope and evidence for customers to decide: region, zone, data plane, control plane, identity path, API family, feature, or dependency. GitHub's 2018 incident is a useful reminder that data integrity may matter more than immediate restoration; GitHub chose extended degradation rather than risk user data consistency \[19\].

Pause guidance is often missing. Customers may need to stop imports, scheduled jobs, billing runs, migrations, or repeated manual actions. "Degraded performance" does not tell them whether the risk is delay, loss, duplicate processing, stale reads, failed writes, or inconsistent exports.

Communication guidance also matters because enterprise customers are often providers themselves. Atlassian's incident templates are useful because they separate investigating, identified, monitoring, and resolved phases, and because "next update" is not the same as "fix complete" \[15\].

Waiting can be the right action when it is supported by evidence. "Monitoring" should point to recovery signals: errors below threshold, queues draining, synthetic checks passing, control-plane APIs restored, or customer-specific resource state returning to available.

## 5. A customer reliability record

A customer reliability record is a structured statement of service health intended to support customer decisions. It can appear as a page, API object, support attachment, incident update, or post-incident record. The essential fields are:

| Field | Purpose |
| --- | --- |
| Identity | Stable ID, version, owner, created time, and updated time |
| Relevance | Public, account-specific, tenant-specific, resource-specific, or unknown |
| State | Investigating, identified, mitigating, monitoring, resolved, or closed |
| Impact | Affected customer capabilities in plain language |
| Scope | Products, regions, zones, tenants, resources, API methods, cohorts, and time window |
| Confidence | Observed, probable, possible, not observed, or unknown, with evidence source |
| Dependency | Control plane, data plane, identity, monitoring, upstream provider, or customer configuration |
| Customer action | Retry, fail over, pause, degrade, communicate, wait, or contact support |
| Data and security | Confidentiality, integrity, availability, and privacy impact where known |
| Recovery evidence | What has recovered, what has not, and which signals support the state |
| Next update | Time or condition for the next update |
| Residual risk | Backlogs, reconciliation, recurrence risk, or affected edge cases |

The record is a translation layer. It does not reveal internal dashboards. It maps internal knowledge to customer choices.

State and decision must remain separate. "Investigating" may mean customers should avoid repeated manual retries. "Identified" may mean only customers in one region should consider failover. "Monitoring" may mean low-risk work can resume while irreversible jobs remain paused. "Resolved" may mean customers should reconcile activity in the affected time window.

Confidence should also be explicit. A practical model is enough:

- Observed: provider telemetry or resource health shows impact.
- Corroborated: independent signals agree.
- Probable: the customer uses an affected service, but direct evidence is incomplete.
- Possible: a shared dependency is implicated.
- Not observed: no evidence within stated monitoring limits.
- Unknown: the relevant signal is unavailable or outside provider visibility.

Many customers need this as an API event. A reliability API should expose stable values for state, severity, confidence, scope, and recommended actions. Statuspage's API already contains concepts such as components, incident updates, impact, status, timestamps, and postmortem fields \[16\]. Provider-specific health APIs go further by making relevance account-specific or project-specific \[9\]\[10\].

## 6. Status pages, APIs, and support channels

A public status page is a necessary floor. It is not the ceiling.

Public pages serve several functions:

- they confirm that the provider recognises an issue;
- they give unauthenticated access when login is impaired;
- they create a public incident history;
- they reduce duplicate support contacts;
- they give customers language for their own stakeholders.

But public pages are deliberately coarse. They cannot list every affected tenant. They should not expose sensitive implementation details. They may be unavailable or delayed if the provider's communication tools share dependencies with the affected service.

AWS's December 2021 event is a clear public example. The same event that affected multiple AWS services also delayed internal monitoring, impaired CloudWatch data, affected support case creation, and prevented Service Health Dashboard tooling from failing over as intended. AWS later said it expected a new version of the Service Health Dashboard to make impact easier to understand and a more resilient support system architecture \[18\].

The point is not that AWS is uniquely weak. Communication systems are reliability systems in their own right: their dependencies, failover paths, permissions, templates, and staffing should be designed and tested as part of the product.

### Use multiple channels with one source of truth

Customers need the same incident facts through different channels:

- public status page for broad awareness;
- authenticated dashboard for customer-specific relevance;
- API and event stream for automation;
- email, SMS, Slack, Teams, or webhook for subscribers;
- support case context for account-specific evidence;
- post-incident report for learning and contractual review.

Atlassian Statuspage materials describe components, subscriber notifications, incident statuses, templates, postmortems, and APIs \[14\]\[15\]\[16\]\[17\]. The important design point is less about the vendor than about structuring communication operationally before the incident happens.

### Keep the public fallback independent

Authenticated health is valuable because it is relevant. It also has a dependency problem: a customer may need identity, console access, or API access to see the personalised message. Google Cloud explicitly recommends using the public Cloud Service Health dashboard and RSS feed as fallback channels if Personalized Service Health is unavailable or inaccessible \[11\].

Every provider should test the customer question: "If the service is broken enough that I cannot sign in, where do I get trustworthy information?"

### Support should not be the only path

Support teams are essential during enterprise incidents. They are not a scalable substitute for a health surface. If every customer has to open a ticket to ask "am I affected?", the provider has created a communication denial of service against itself.

The better pattern is layered:

1. public acknowledgement for broad events;
2. authenticated relevance for affected customers;
3. API integration for customer operations;
4. support escalation for exceptions, regulated obligations, or ambiguous evidence.

## 7. Dependencies, uncertainty, and recovery evidence

Reliability information becomes decision-grade when it shows the shape of uncertainty.

### Dependencies

Dependencies matter because the customer's mitigation often depends on which layer failed. A control-plane incident may block provisioning while existing workloads continue. A data-plane incident may affect live requests. An identity incident may block access to the console and to the health dashboard. A monitoring incident may reduce the provider's confidence in every other statement.

Public postmortems show these patterns repeatedly. AWS described an internal network congestion event that affected foundational services, monitoring, DNS, authorisation-related paths, control planes, and customer-facing communication tools \[18\]. Slack described a network issue in a cloud dependency that impaired Slack, while its own dashboarding and alerting service became unavailable and made debugging harder \[22\]. GitHub described a short network partition that triggered database topology changes and led to a long recovery because of data consistency constraints \[19\].

A customer reliability record should therefore identify dependency class without exposing unnecessary topology:

- provider control plane;
- provider data plane;
- identity and access;
- monitoring and health reporting;
- third-party provider;
- customer configuration;
- regional network path;
- data replication or backlog processing.

Customers do not always need names of internal services. They do need to know whether an affected dependency changes their safe action.

### Uncertainty

Incident communication often begins before root cause is known. Waiting for certainty can leave customers with no usable information. Publishing guesses as fact can destroy trust.

The record should use disciplined language:

| Phrase | Meaning |
| --- | --- |
| We are investigating reports of... | Symptoms are known, root cause is not confirmed |
| We have identified impact to... | Affected capability and scope are confirmed enough to guide action |
| We believe... | Best current explanation, still subject to change |
| We have not observed... | No evidence within stated monitoring limits |
| We do not yet know... | Evidence is missing or inconclusive |
| We have mitigated... | Immediate customer impact reduced, root cause may remain |
| We have resolved... | Recovery criteria met for stated scope |

A promise of the next update time is often more valuable than a premature recovery time. Atlassian's incident communication templates explicitly separate "next update" from "fix complete" \[15\].

### Recovery evidence

A resolved update should answer "why do you believe this is over?"

Useful recovery evidence can include:

- error rates below threshold for a stated interval;
- latency within objective for affected flows;
- backlog drained or draining at a measured rate;
- queued events replayed or explicitly declared unrecoverable;
- synthetic checks passing from relevant locations;
- customer-specific resource state returning to available;
- failed control-plane operations retried successfully;
- reconciliation completed for affected writes;
- support-contact patterns returning to baseline;
- post-mitigation canary or rollback validation.

Fastly's public account is strong because it gives times: detection within one minute, status post at 09:58 UTC, impacted services beginning to recover at 10:36 UTC, majority recovered at 11:00 UTC, and incident mitigated at 12:35 UTC \[20\]. GitHub's public account is strong because it explains the recovery constraint as well as the duration: protecting data consistency took priority over faster restoration \[19\].

Recovery evidence should not become a data dump. It should state the reason customers can change behaviour.

## 8. Why raw telemetry and excessive transparency can harm

The cure for vague reliability communication is not to publish every graph, log line, trace, and internal chat message.

Raw telemetry can harm customers in several ways.

First, it can expose sensitive data. Logs and traces often include identifiers, file paths, internal hostnames, request bodies, headers, tokens, customer names, and other sensitive material. OWASP advises that sensitive personal data, access tokens, passwords, database connection strings, encryption keys, higher-classification data, and commercially sensitive information should usually not be recorded directly in logs, or should be removed, masked, sanitised, hashed, or encrypted \[30\]. NIST's log-management guidance exists because collecting, protecting, analysing, and retaining logs is itself a security and governance problem \[31\].

Second, it can help attackers. During a security-related incident, detailed topology, mitigations in progress, unpatched weaknesses, or detection gaps can help an attacker adapt. Real-time communication therefore needs a security review path and pre-approved safe language.

Third, it can overwhelm decision makers. Raw metrics are optimised for engineers who know the system. Customers need a decision. Ten charts may show that an incident is real without telling a customer whether writes are safe. OWASP's warning that logging should match its intended purpose applies equally to reliability communication \[30\].

Fourth, it can create false precision. Server-side success rate does not include requests that never arrived. An average hides tail latency. A resource health state may lag a symptom. If the provider publishes a number without scope and confidence, customers may treat it as stronger than it is.

Finally, it can trigger unsafe action. Customers who see high error rates without context may restart clients, drain queues, fail over databases, or disable protections in ways that worsen recovery.

The target is usable candour rather than total transparency: enough information to act, with enough restraint to protect customers.

## 9. Operating model for enterprise leaders

Customer-usable reliability is not owned by one team. Product leaders define the customer decisions the surface must support. Reliability engineers define signals, error budgets, dependencies, and recovery evidence. Customer engineering and support teams translate impact into workflows customers recognise. Security and legal teams define what cannot be exposed, what must be disclosed, and when regulated notifications are required.

A practical operating model has five layers.

### 1. Define customer-critical journeys

Begin with what customers do, not with components. Examples include sign in, checkout, receive event, provision environment, export report, reconcile payment, or view audit history. For each journey, define success criteria, acceptable delay, data-integrity requirement, fallback path, and owner. Azure's Well-Architected guidance recommends scoring user and system flows by business importance and using those scores for design, testing, and incident management \[6\].

### 2. Map internal signals to customer outcomes

For each journey, map internal signals to customer-visible states and blind spots.

| Internal signal | Customer interpretation |
| --- | --- |
| API 5xx rate by method | Requests may fail. Retry guidance depends on idempotency. |
| Queue age | Events may be delayed but not lost if retention is intact. |
| Replica lag | Reads may be stale. Writes may still be durable. |
| Control-plane error | Existing workloads may run while changes fail. |
| Identity errors | Users may be unable to sign in or rotate credentials. |
| Monitoring delay | Provider confidence is reduced. Customer monitoring matters more. |

### 3. Build the reliability surface

The surface should include a public page, authenticated customer impact records, machine-readable events, support tooling that uses the same record, templates for uncertain states, accessible summaries, and post-incident reports linked to the original timeline. It should be versioned so customers can see what was known, when it changed, and what evidence supported closure.

### 4. Exercise the communication path

Test cases should include identity unavailable, status tooling unavailable, support systems impaired, telemetry delayed, third-party dependency at fault, security incident with limited disclosure, partial recovery with backlog still draining, and customers consuming API events rather than web pages. CISA's federal playbooks emphasise standard procedures to identify, coordinate, remediate, recover, and track successful mitigations \[24\]. Customer communication should be part of that rehearsal.

### 5. Measure whether customers used it

Useful measures include time to first customer-usable update, percentage of updates with scope and confidence, support cases asking questions already answered by the record, API events routed into customer runbooks, correction rate for impact scope, time from mitigation to recovery evidence, accessibility conformance, and post-incident customer feedback.

An absence of complaints is not proof that the record worked. Customers may have built their own workarounds, stopped trusting the page, or suffered quietly.

## 10. Strongest counterargument

The strongest counterargument is that customer-specific reliability reporting can make reliability worse.

A provider cannot fully understand every customer's architecture. Low-confidence guidance can trigger unnecessary failover, duplicate processing, or avoidable damage inside customer systems. Detailed disclosure can expose internal systems, customer relationships, security controls, or regulated data. Heavy legal review can make updates slower. A cautious public status page may seem safer than a complex personalised system that creates false confidence.

This argument should be taken seriously. A provider is not the customer's SRE team. It should not tell a customer to perform a risky action when it cannot know the customer's architecture. It should not publish security-sensitive detail merely to look transparent.

But the counterargument does not support silence. Customers already make decisions during provider incidents. With vague status, they infer scope, cause, and recovery from fragments: social media, support queues, their own alarms, and other customers' reports. Silence does not prevent reliance. It makes reliance less informed.

The right response is a bounded model:

- distinguish provider evidence from customer responsibility;
- state confidence and detection limits;
- give decision categories rather than commands where context is missing;
- expose account-specific relevance only where evidence supports it;
- redact sensitive detail by design;
- keep raw telemetry behind appropriate controls;
- preserve customer monitoring as a required control.

Customer reliability information should not say "you are safe" unless the provider can support that claim. It can say "we have observed impact to these resources", "we have not observed impact within these limits", "writes may have succeeded without acknowledgement", or "do not retry non-idempotent requests".

That is operationally useful honesty rather than overreach.

## 11. Limitations and source quality

This paper is based on public information available through 25 August 2026. It uses vendor documentation, public post-incident reports, standards, regulatory material, and public service-health documentation. It does not use private employer information, non-public incident data, customer records, or confidential metrics.

The strongest sources are primary documents: SRE books and workbooks from Google, cloud-provider service-health documentation, official public postmortems, NIST publications, CISA playbooks, W3C accessibility standards, and official EU regulatory materials. Vendor documents have incentives: they may emphasise progress, omit internal disagreement, and describe product capabilities more than empirical outcomes. Public postmortems are selected records, not a complete incident population. Regulations define obligations for particular sectors and jurisdictions, not universal product design.

Cloud well-architected guidance from AWS and Google is used as architecture context, not as proof of customer-facing reporting quality \[5\]\[7\]. Regulatory sources are used only to frame obligations, not to define universal product design: NIST incident response and CSF 2.0, DORA, and NIS2 emphasise incident handling, recovery, communication, or reporting in specific governance contexts \[23\]\[25\]\[26\]\[27\].

The evidence supports these narrower conclusions:

1. SLOs and error budgets are accepted reliability decision tools inside engineering organisations \[1\]\[2\]\[3\]\[4\].
2. Major providers already expose increasingly personalised health information through dashboards, APIs, logs, and alerts \[8\]\[9\]\[10\]\[11\]\[12\]\[13\].
3. Public incidents show that dependency failures, monitoring loss, delayed communication, and recovery uncertainty are real operating conditions \[18\]\[19\]\[20\]\[21\]\[22\].
4. Security and privacy guidance limits what can safely be exposed as raw telemetry \[30\]\[31\].
5. Accessibility and plain language matter because urgent reliability information must be usable under stress and by people using different tools \[28\]\[29\].

The evidence does not prove that every customer would benefit from every field in the proposed record. It does not quantify the market value of better reliability reporting. It does not show a complete public benchmark comparing vendor health surfaces. Those are open areas for research and product validation.

## 12. What this paper does not claim

This paper does not claim that internal SLOs are obsolete. They remain one of the best tools for engineering trade-offs.

It does not claim that customers should receive raw dashboards, internal logs, stack traces, root-cause theories, or security-sensitive topology.

It does not claim that a provider can know the full business impact of an incident inside every customer environment.

It does not claim that a status page can replace a customer's own monitoring, runbooks, contract review, or business continuity plan.

It does not claim that transparency means immediate disclosure of every technical detail. During security incidents, excessive detail can harm customers.

It does not claim that regulation alone will produce useful reliability communication. Regulatory notifications and customer operational guidance have different audiences and time scales.

It does not claim that every outage requires a long postmortem. The amount of documentation should match customer impact and learning value.

The claim is this:

Customers need reliability information in a form they can act on. Internal health is an input. The product surface should translate that input into scoped, evidence-based, accessible, and secure guidance for customer decisions.

## Conclusion

Reliability is often discussed as an engineering property. A service is up or down, within SLO or outside it, burning budget slowly or quickly. Those views are necessary. They are also incomplete.

For a customer, reliability is experienced as the ability to continue useful work with acceptable risk. That experience depends on the provider's systems, the customer's architecture, shared dependencies, data integrity, recovery time, and the information available during uncertainty.

A customer who receives only a red component still has to guess. A customer who receives a decision-grade reliability record can choose. They can retry safely, fail over deliberately, pause risky work, communicate honestly, reconcile affected data, or wait with evidence.

The shift is not from secrecy to radical transparency.

It is from provider-centred health to customer-centred reliability.

That shift requires product work. It needs schemas, APIs, accessible pages, incident templates, security review, support integration, and post-incident learning. It requires product managers, SREs, customer engineers, support leaders, security teams, legal teams, and executives to agree that communication is part of the reliability system.

The useful question is not whether the provider knows every internal metric.

It is whether the customer can make a better decision because of what the provider chose to share.



## About the author

Jason Doyle writes about reliable software, observability, applied AI, and practical controls for systems that influence human decisions. He publishes at [jasondoyle.ie](https://jasondoyle.ie) and can be contacted at [contact@jasondoyle.ie](mailto:contact@jasondoyle.ie).

## References

1. Google SRE, _Service Level Objectives_, Site Reliability Engineering, <https://sre.google/sre-book/service-level-objectives/>.
2. Steven Thurgood and David Ferguson, _Implementing SLOs_, Google SRE Workbook, <https://sre.google/workbook/implementing-slos/>.
3. Steven Thurgood, _Example Error Budget Policy_, Google SRE Workbook, 19 February 2018, <https://sre.google/workbook/error-budget-policy/>.
4. Steven Thurgood et al., _Alerting on SLOs_, Google SRE Workbook, <https://sre.google/workbook/alerting-on-slos/>.
5. Amazon Web Services, _Reliability Pillar - AWS Well-Architected Framework_, 6 November 2024, <https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html>.
6. Microsoft Learn, _Architecture strategies for defining reliability targets_, Azure Well-Architected Framework, archived 10 February 2026, <https://web.archive.org/web/20260210140312/https://learn.microsoft.com/en-us/azure/well-architected/reliability/metrics>.
7. Google Cloud, _Reliability pillar in the Google Cloud Well-Architected Framework_, last reviewed 30 December 2024, <https://cloud.google.com/architecture/framework/reliability>.
8. Amazon Web Services, _What is AWS Health?_, AWS Health User Guide, <https://docs.aws.amazon.com/health/latest/ug/what-is-aws-health.html>.
9. Amazon Web Services, _AWS Health API Reference_, archived 15 February 2025; source states last published 14 February 2025, <https://web.archive.org/web/20250215015708/https://docs.aws.amazon.com/health/latest/APIReference/Welcome.html>.
10. Google Cloud, _Personalized Service Health overview_, <https://cloud.google.com/service-health/docs/overview>.
11. Google Cloud, _Google Cloud incident communication_, <https://cloud.google.com/service-health/docs/incident-communication>.
12. Microsoft Learn, _What is Azure Service Health?_, updated 31 October 2025, <https://learn.microsoft.com/en-us/azure/service-health/overview>.
13. Microsoft Learn, _Azure Resource Health overview_, updated 26 February 2026, <https://learn.microsoft.com/en-us/azure/service-health/resource-health-overview>.
14. Atlassian Support, _Incident communication tips_, <https://support.atlassian.com/statuspage/docs/incident-communication-tips/>.
15. Atlassian, _Learn incident communication with Statuspage_, <https://www.atlassian.com/incident-management/tutorials/incident-communication>.
16. Atlassian, _Statuspage API documentation_, <https://developer.statuspage.io/>.
17. Atlassian Support, _Create a postmortem_, <https://support.atlassian.com/statuspage/docs/create-a-postmortem/>.
18. Amazon Web Services, _Summary of the AWS Service Event in the Northern Virginia (US-EAST-1) Region_, 10 December 2021, <https://aws.amazon.com/message/12721/>.
19. GitHub, _October 21 post-incident analysis_, 30 October 2018, <https://github.blog/news-insights/company-news/oct21-post-incident-analysis/>.
20. Fastly, _Summary of June 8 outage_, 8 June 2021, <https://www.fastly.com/blog/summary-of-june-8-outage>.
21. Cloudflare, _Details of the Cloudflare outage on July 2, 2019_, 12 July 2019, <https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019/>.
22. Slack Engineering, _Slack's outage on January 4th, 2021_, 26 January 2021, <https://slack.engineering/slacks-outage-on-january-4th-2021/>.
23. National Institute of Standards and Technology, _Computer Security Incident Handling Guide_, SP 800-61 Rev. 2, August 2012, <https://csrc.nist.gov/pubs/sp/800/61/r2/final>.
24. Cybersecurity and Infrastructure Security Agency, _Federal Government Cybersecurity Incident and Vulnerability Response Playbooks_, November 2021, <https://www.cisa.gov/resources-tools/resources/federal-government-cybersecurity-incident-and-vulnerability-response-playbooks>.
25. National Institute of Standards and Technology, _The NIST Cybersecurity Framework (CSF) 2.0_, NIST CSWP 29, 2024, <https://www.nist.gov/publications/nist-cybersecurity-framework-csf-20>.
26. European Union, _Regulation (EU) 2022/2554 on digital operational resilience for the financial sector_, Official Journal, 2022, <https://eur-lex.europa.eu/eli/reg/2022/2554/oj/eng>.
27. European Commission, _The NIS2 Directive_, <https://digital-strategy.ec.europa.eu/en/policies/nis2-directive>.
28. World Wide Web Consortium, _Web Content Accessibility Guidelines (WCAG) 2.2_, 5 October 2023, <https://www.w3.org/TR/WCAG22/>.
29. US General Services Administration, _Plain language_, <https://digital.gov/guides/plain-language>.
30. OWASP Cheat Sheet Series, _Logging Cheat Sheet_, <https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html>.
31. National Institute of Standards and Technology, _Guide to Computer Security Log Management_, SP 800-92, September 2006, <https://csrc.nist.gov/pubs/sp/800/92/final>.
