Observability strategy

Observability Is a Decision System

Why telemetry only matters when it changes a decision.

Jason Doyle First published: 1 May 2026 27 minute read

Disclosure: These views are my own and do not represent my current or any former employers.

Executive summary

Telemetry is inventory. Logs, metrics, traces, profiles, events, and spans are records of things that happened or are happening. They are necessary, but they are not the value of observability.

Observability creates value when evidence reaches a person or system that can make a better decision. That decision may be made by an on-call engineer, an incident commander, a release controller, a product owner, a customer communications lead, a support organisation, or an executive deciding where to invest. If the evidence does not change detection, interpretation, action, communication, prioritisation, or learning, it is stock on a shelf.

This paper argues for starting observability architecture with the decision rather than with the signal type. Control theory used observability to describe whether internal state can be inferred from outputs [1]. OpenTelemetry describes observability as the ability to understand a system from the outside, ask questions without knowing the internals in advance, and handle unknown unknowns; it also defines telemetry signals as outputs such as traces, metrics, logs, baggage, and profiles [2][3]. Both traditions point to the same constraint: outputs matter because someone needs to infer state and choose what to do.

SRE literature is already decision-centred when read closely. Google's SRE book says pages should interrupt people only for real problems that need investigation and mitigation, and warns that frequent pages cause people to skim or ignore noise [4]. Google's SLO guidance treats objectives and error budgets as tools for deciding reliability work, feature velocity, and customer expectations [5]. Its SLO alerting chapter evaluates alerting strategies by precision, recall, detection time, and reset time [6].

DORA research also points away from tool worship. It defines monitoring as predefined metrics or logs and observability as the ability to debug properties and patterns not defined in advance [7]. Its delivery metrics are meant to help teams assess performance, prioritise improvement, and validate progress, while warning against gaming metrics and measurement at the expense of improvement [8]. The 2024 report emphasised stable priorities and learning; DORA also frames user needs and feedback loops as measurable [9][10].

The technical standards and operating practices matter because decision systems need evidence that can travel and people who can act on it. Trace context, semantic conventions, distributed tracing, troubleshooting, incident response, and postmortems all contribute to that loop [11][12][13][14][15][16].

Public incident reports show the cost of weak decision coverage. Dashboard dependencies, broken internal tools, dropped logs, missing customer contacts, and untested recovery paths all affect what responders can decide when the system is under stress [17][18][19][20][21].

This paper proposes a measurable decision coverage model for Principal and Director-level leaders. For every material operational decision, define the decision, owner, evidence, interpretation rule, action path, feedback loop, and validation method. Score each element from 0 to 1, weight the decision by consequence, and use the weakest element as the conservative coverage score. Track decision latency, evidence quality, false-confidence gaps, and feedback closure alongside ingestion volume and cost.

The aim is not to replace dashboards with forms. It is to ensure that dashboards, alerts, traces, and logs have a job. Observability should answer: who can decide, what evidence do they need, how fast must it arrive, what action can they take, and how will the system learn whether that decision improved the next one?

1. Telemetry is not the outcome

A modern production system can emit more data than any team can read. That data has operational value only when it becomes evidence.

Telemetry is emitted by systems. Evidence is telemetry that has been selected, preserved, correlated, and presented for a judgement. A log line saying that a request failed is telemetry. A view showing that checkout failures rose for new customers in one region after a payment-provider routing change is evidence. The difference determines cost and architecture.

If the goal is telemetry inventory, the programme optimises for coverage of signal types, retention, indexing, sampling, dashboards, and vendor integration. These are real concerns. They also tend to generate the question, "What should we collect?"

If the goal is decision quality, the programme begins with different questions:

  • What decision must be made?
  • Who or what owns it?
  • What evidence would change the decision?
  • How soon must that evidence arrive?
  • How will the owner interpret it?
  • What action is available?
  • How will we know whether the decision worked?

This does not make telemetry less important. It makes telemetry accountable.

A service may have excellent metric coverage and still fail a decision. The dashboard may be green because it measures infrastructure rather than the user journey. The trace may exist but omit the tenant, region, feature flag, or operation that would reveal the affected cohort. The alert may fire but page a team that cannot act. In each case, telemetry exists. Observability failed.

A small set of well-chosen signals can do better when they map to clear ownership and action. A synthetic checkout probe can tell an incident commander that revenue-impacting functionality is failing. A burn-rate alert can tell an on-call engineer that an SLO is being consumed fast enough to require response. A deployment event linked to error increase can tell a release system to halt progression.

Observability is therefore not the amount of data a system emits. It is the degree to which the organisation can turn system behaviour into timely, justified action.

2. Begin with the decision

The smallest useful unit of observability design is a decision, not a dashboard.

A decision can be human or automated. A human may decide to declare an incident, roll back a release, disable a feature, inform customers, pause a launch, approve capacity spend, change an SLO, or invest in resilience work. A system may decide to page, autoscale, shed load, open a ticket, block a deployment, or route traffic away from a region.

The decision should be written in operational language:

Weak form Decision form
We need database dashboards. Should the on-call engineer fail over, reduce load, or keep investigating?
We need traces. Which service boundary is adding latency for this user journey?
We need executive reliability metrics. Should the leadership team invest in resilience, accept risk, or change product promises?
We need better alerts. Which condition should interrupt a named owner within five minutes because action can prevent budget exhaustion?

The decision form changes instrumentation. It forces the designer to ask whether the evidence is sufficient for the action.

For example, a checkout SLO does not need every internal metric on the first page. It needs to show whether users can complete checkout, whether the failure is growing, which cohorts are affected, what changed recently, and what mitigation is available. If rollback is the likely action, deployment events and feature flags are first-class evidence. If communication is the likely action, affected tenant, geography, time window, and workaround are first-class evidence.

This approach also changes ownership. A dashboard cannot own a decision. A team, role, or automated controller must own it, with authority to act. A page to someone who cannot change traffic, roll back, disable a feature, contact the vendor, or escalate to incident command is delay with a ringtone, not decision support.

Google's incident response guidance distinguishes resolving an incident from managing one, and emphasises command, communication, control, defined roles, and a working record [15]. Observability architecture should support those roles explicitly. The incident commander needs impact, scope, trajectory, confidence, and decisions made. The operations lead needs hypotheses, safe mitigations, and evidence of effect. The communications lead needs customer-visible facts, affected populations, uncertainty, and next update times. A single technical dashboard rarely serves all three.

3. Decision coverage

Decision coverage is a way to measure whether observability supports the decisions the organisation actually needs to make.

The model starts with a decision register. Each row describes one decision that matters to reliability, customer trust, engineering flow, security, cost, or compliance. It should focus on consequential decisions that repeat, have high impact, or create confusion during incidents.

Each decision has seven coverage elements:

Element Question Score
Purpose Is the decision stated in terms of user, business, safety, security, or operational consequence? 0, 0.5, 1
Owner Is there a named owner or automated controller with backup coverage? 0, 0.5, 1
Evidence Is the required evidence available with defined freshness, scope, and quality? 0, 0.5, 1
Interpretation Is there a rule, playbook, SLO, hypothesis path, or trained judgement for reading the evidence? 0, 0.5, 1
Action Can the owner take or trigger a concrete action inside the required time? 0, 0.5, 1
Feedback Is the outcome reviewed and used to improve signals, playbooks, thresholds, or ownership? 0, 0.5, 1
Validation Has the decision path been tested by drill, game day, incident review, replay, or sampling? 0, 0.5, 1

Use 0 when absent, 0.5 when informal or partial, and 1 when explicit, current, and tested.

A conservative score uses the weakest element because the weakest element is where the decision fails:

Decision coverage for one decision = weight x min(Purpose, Owner, Evidence, Interpretation, Action, Feedback, Validation)
Portfolio decision coverage = sum(decision coverage) / sum(decision weights)

Weight should reflect consequence, not visibility. A low-volume payment failure, data-loss risk, or security containment decision can carry more weight than a common infrastructure warning. A practical scale is 1 for convenience, 2 for operational effect, 3 for customer-visible effect, 4 for material contractual or financial effect, and 5 for safety, security, legal, or existential risk.

The model also calculates a false-confidence gap:

Support score = average of the seven element scores, weighted by decision consequence
False-confidence gap = support score - conservative coverage score

A high support score with low conservative coverage means the organisation has many visible pieces but at least one broken link. There may be dashboards but no owner, an alert but no authority, or an SLO but no feedback loop when the SLI fails to match customer pain.

Decision coverage should be reviewed by service, user journey, and executive objective. Principal engineers can use it to identify architectural and instrumentation gaps. Directors can use it to decide whether reliability investment is reducing the right operational risks rather than simply expanding the telemetry bill.

4. User journeys, SLOs, and error budgets

Observability should be anchored in user journeys because users experience services as outcomes, not components. A customer trying to sign in, search, pay, upload, invite a colleague, receive a notification, or recover an account cares whether the task succeeded, whether it was fast enough, whether data was correct, and whether failure left a clear path.

SLOs translate part of that experience into a decision mechanism. Google's SLO guidance defines an SLI as a quantitative measure of service level and an SLO as a target for that measure. It recommends SLIs as ratios of good events to total events and treats error budgets as a way to make reliability tradeoffs explicit [5]. The workbook also warns that without organisational commitment to use the error budget for decision making, SLO compliance becomes just another KPI [5].

That warning is central to this paper.

An SLO should not be a chart that proves the team is good. It should be a policy that decides something. If a service burns error budget rapidly, the team may slow releases, roll back, shed load, add capacity, escalate dependency risk, or inform customers. If an SLI does not correlate with support, revenue, retention, or user research signals, the organisation should refine it.

SLO alerting is a decision-latency problem. Google's SLO alerting chapter uses precision, recall, detection time, and reset time to compare alerting strategies [6]. These are exactly the qualities a decision system needs. Low precision pages people for events that do not matter. Low recall misses events that do matter. Slow detection consumes the budget before anyone can act. Slow reset leaves responders unsure whether mitigation worked.

User journeys also reveal where classic infrastructure metrics are insufficient. CPU, memory, queue depth, database connections, and garbage collection pauses are useful. They do not by themselves say whether a user could complete checkout. A journey-level SLO may use synthetic monitoring, server-side outcomes, client-side measurements, and support reports. Google's SLO guidance distinguishes an SLI specification from implementations that vary in quality, coverage, and cost [5].

Decision coverage extends this by asking: for each important journey, what decisions rely on the SLO?

  • Does the on-call engineer know when to page another team?
  • Does the release system know when to halt rollout?
  • Does the product owner know when reliability risk should outrank feature work?
  • Does customer support know who is affected and what workaround exists?
  • Does leadership know whether reliability investment is buying down user-visible risk?

A journey without these decision links is being measured, not managed.

5. Instrumentation design as evidence design

Instrumentation should be designed backwards from the evidence a decision requires. This means building common, structured evidence that can answer recurring decisions without fresh code changes during every incident, rather than adding bespoke telemetry for every meeting question.

Good evidence has at least eight qualities:

Quality Meaning
Fidelity It measures the thing the decision cares about, preferably from the user's point of view.
Coverage It includes the relevant journeys, tenants, regions, clients, jobs, and dependencies.
Granularity It can separate the affected cohort from the healthy majority.
Correlation It links requests, logs, metrics, deployments, feature flags, and owners.
Freshness It arrives quickly enough for the decision.
Integrity It is trustworthy, retained, and not silently dropped in the failure mode under review.
Cost control It keeps useful detail without unlimited storage or query cost.
Interpretability It uses names, units, attributes, and semantics that responders understand.

OpenTelemetry is useful here because it standardises signals and context rather than treating every library and platform as a private language. Its semantic conventions define common names for operations and data across traces, metrics, logs, profiles, and resources [12]. W3C Trace Context standardises request context propagation across services [11]. OpenTelemetry's signal model includes trace, metric, log, and baggage data that can be correlated when instrumentation preserves the right fields [3].

Standardisation is not sufficient. A trace with no user journey, tenant class, feature flag, deployment, or error semantics may show a path while hiding the decision. A high-cardinality metric may be powerful for debugging but unavailable at the retention window needed for review. A log field may be present but inconsistent across languages.

A decision-centred instrumentation review asks questions such as:

  • Which attribute would let us distinguish a premium customer journey from a background retry?
  • Which span boundary marks the moment value was or was not delivered to the user?
  • Which event records the configuration or release that changed behaviour?
  • Which evidence will still exist if the observability backend is degraded?
  • Which fields must be redacted, aggregated, or separated for privacy and security?
  • Which sampling policy preserves rare but high-consequence failures?

The answer may be a trace, a metric, a log, an event, a profile, a synthetic probe, a ticket field, or a customer-support signal. The important point is that the evidence is selected because of a decision, not because a tool category was missing from a slide.

6. Unknown unknowns and exploratory debugging

A decision-first approach must not become a closed list of known failures. The best argument for observability, rather than monitoring alone, is that modern systems fail in combinations that were not predicted.

DORA distinguishes monitoring as predefined metrics or logs from observability as the ability to debug properties and patterns not defined in advance [7]. OpenTelemetry's primer makes a similar claim: observability helps troubleshoot novel problems and ask why something is happening [2]. Dapper's history at Google is also instructive. Its designers did not anticipate every analysis tool that later used the trace data; the value came from a tracing substrate that enabled developers and operations teams to ask new questions [13].

Exploration still has decisions inside it. Google's troubleshooting chapter describes troubleshooting as an iterative process of forming hypotheses, testing them against observations, and taking corrective action [14]. The decision goes beyond "what is broken?" to include "which hypothesis should I test next?", "is this evidence sufficient to mitigate?", and "should I stop root-causing and first make the system safer?" The chapter explicitly recommends stopping the bleeding before full root-cause analysis in major incidents [14].

The implication is practical. Observability platforms should support both predetermined decision paths and exploratory investigation. Predetermined paths need durable, simple indicators: SLO burn, journey success, queue saturation, synthetic checks, deployment changes, backup success, certificate expiry, and security-relevant events. Exploratory paths need flexible slicing, high-cardinality attributes where justified, request correlation, recent change history, and enough raw or structured detail to test unexpected hypotheses.

These two modes should not compete. The incident starts with known decision support: is the user journey failing, who owns response, what mitigation is safe? Exploration then narrows the cause. The postmortem feeds learning back into both layers. Some exploratory questions become new routine checks. Some routine alerts are retired because they did not lead to useful action.

7. Incidents expose decision architecture

Incidents are where observability architecture becomes visible. During normal operation, a weak observability system may appear adequate. The problem appears when time pressure, dependency failure, customer impact, and incomplete knowledge collide.

Slack's January 2021 outage shows observability dependency as operational risk. External monitoring first paged the team, but dashboarding and alerting then became unavailable. Engineers still had consoles, status pages, command-line tools, logs, and direct metric queries, but debugging was less efficient. Autoscaling also deprovisioned instances that responders were using for investigation [17]. Dashboards were not the failure. The absence of fallback evidence and a responder-safe operating mode was.

Meta's October 2021 outage shows a different impairment. The network event disconnected data centres, DNS became unreachable, and normal internal tools used for investigation and recovery were broken. Engineers had to be sent onsite, and physical and security controls added recovery time [18]. Recovery decisions depend on access paths as much as on telemetry content.

Cloudflare's July 2020 outage demonstrates the importance of understanding what evidence may be missing. After a configuration error caused traffic disruption and recovery, a separate period of core congestion caused some logs to be dropped while the edge continued operating [19]. A post-incident analysis that assumes logs are complete would be more confident than the evidence allows.

Atlassian's April 2022 outage shows that decision coverage extends beyond engineering diagnosis. The company reported that customer contact information for some affected customers was deleted, complicating direct communication, and that restoration for many customers in a live environment required complex, partly manual work [20]. Observability for such an incident includes contact paths, restoration status, validation state, support workflow, and executive communication.

GitLab's 2017 database incident shows why evidence quality and runbook clarity matter before disaster. Its postmortem reported that multiple backup and replication approaches were not working reliably or were not properly set up, and that documentation gaps contributed to confusion [21]. The decision "restore from backup" is covered only if existence, freshness, restore process, validation, ownership, and expected data loss are known before the deletion.

These reports should not be over-generalised. Each system and incident was different. Together, they show that incident outcomes depend on decision architecture: who knew what, when, through which evidence, with what authority, and with what fallback when the normal path failed.

This matches NIST's incident handling lifecycle [22].

8. Dashboard and alert failure modes

Dashboards and alerts fail in predictable ways when they are not tied to decisions.

A dashboard can be true and useless

A chart can accurately show CPU, latency, or queue depth while failing to answer the owner's question. The on-call engineer may need to know whether to roll back. The dashboard may show that errors rose without showing what changed, whether rollback is safe, or which customers are affected.

A dashboard can be green while users are unhappy

Component health does not guarantee journey health. A service can return HTTP 200 while writing the wrong record, omit a notification, serve stale data, or fail only for a small cohort. User-centric SLOs, synthetic checks, client measurements, and support signals help only if they are in the decision path [10].

A dashboard can hide uncertainty

A single availability percentage may hide sampling, delayed ingestion, missing regions, log drops, or excluded traffic. Evidence should show coverage and freshness alongside value. The question is not "what number do we see?" but "what could this number not see?"

An alert can be noisy but still feel safe

Frequent alerts create the appearance of vigilance. They can also teach responders to skim. Google SRE warns that paging humans is expensive and that too many pages lead people to second-guess, skim, or ignore alerts, possibly prolonging outages [4]. Clinical alarm research is not software operations, but it gives a useful warning: high-volume alarms can weaken human response, and targeted alarm management has reduced alarm volume and false alarms in several studies [23].

An alert can be precise and still unowned

A page that detects a condition but reaches the wrong team is a routing failure. A page that reaches the right team but cannot be acted on is an authority failure. A page that requires tribal knowledge is an interpretation failure. Decision coverage treats all three as observability gaps.

A dashboard can become a proxy for accountability

Leaders sometimes ask for dashboards when they need decisions. A reliability dashboard may be used to signal control to executives while no one has agreed what action follows a worsening trend. In that case, the dashboard is theatre. It creates a visible surface without a control loop.

The remedy is to require every recurring dashboard and alert to declare its audience, decision, owner, update frequency, source limits, and retirement condition, rather than simply publishing fewer charts by default.

9. Customer and executive observability

Observability is often designed for engineers first. That is understandable. Engineers instrument systems, receive pages, and debug incidents. But they are not the only decision makers.

Customers need to decide whether to retry, wait, use a workaround, communicate to their own users, invoke contingency plans, or escalate a contract concern. They need accurate status, scope, start time, current state, workaround, next update, and later a credible explanation. They do not need internal metric names unless those names explain impact.

Atlassian's public review after the April 2022 outage is notable because it treated customer communication and contactability as part of the incident, not merely public relations. The review described commitments to earlier public communication, multiple channels, better backup of key contacts, and support tooling for customers who could not use their normal URL or Atlassian ID [20]. That is decision coverage for customers.

Executives need a different view. They should not be asked to read hundreds of traces. They need to decide whether to accept risk, fund reliability work, adjust promises, pause change, change ownership, or communicate materially. The evidence should connect reliability to user impact, error budget consumption, recurrence, cost of delay, architectural risk, and action-item completion.

DORA's metrics guidance warns against using one metric to rule them all, making disparate comparisons, and focusing on measurement at the expense of improvement [8]. That warning is important at executive level. A mainframe batch system, a mobile checkout flow, and an internal AI summarisation tool may require different journeys, tolerances, evidence, and failure definitions.

The useful executive question goes beyond "what is our observability maturity score?" in isolation. It includes:

  • Which material decisions remain uncovered?
  • Which user journeys have no credible SLO?
  • Which incidents are still detected by customers first?
  • Which teams lack authority to act on the alerts they receive?
  • Which postmortem actions remain open because no leader funded them?
  • Which dashboards create confidence that is not supported by evidence quality?

10. Observability maturity

Many maturity models begin with tool adoption: no central logging, central logging, metrics, traces, dashboards, automated alerts, anomaly detection, and perhaps AI-assisted operations. That sequence is understandable because tools are visible and budgetable. It is incomplete.

A decision-centred maturity model has five levels.

Level Description Typical evidence
0: Telemetry fragments Data exists in scattered systems. Decisions rely on memory, direct access, or customer reports. Unowned dashboards, missing runbooks, customer-first detection.
1: Signal inventory Major logs, metrics, traces, and dashboards exist. Alert volume and cost are tracked. Instrumentation catalogue, basic alert routing, central storage.
2: Decision mapping Critical decisions have named owners, evidence, action paths, and latency expectations. Decision register, SLO ownership, incident role views.
3: Validated coverage Decision paths are tested through drills, incident reviews, replay, and sampling. Weak evidence is marked. Coverage scores, false-confidence gap, evidence-quality audits.
4: Learning system Feedback changes instrumentation, automation, training, ownership, and investment priorities. Closed action items, retired alerts, improved decision latency, reduced recurrence.

The jump from level 1 to level 2 is the important one. It is where observability stops being an inventory programme and becomes a management system.

Level 4 is not fully automated. It includes automation where automation is appropriate: canary analysis, release gates, autoscaling, synthetic checks, and policy enforcement. It also includes human learning. Google's postmortem guidance says postmortems are valuable when written well, acted upon, and widely shared, and that action items need clear ownership and tracking [16]. A mature observability programme closes that loop.

Maturity should also be measured by how the system behaves during degradation. If the observability backend is coupled to the service under investigation, the decision system may fail when needed most. If only one expert can interpret a dashboard, coverage is fragile. If logs are sampled away for the cohort that matters, the evidence is weak. If dashboards are not reviewed after incidents, maturity is performative.

11. Practical metrics for a decision system

A decision system needs its own metrics. The following set is small enough to run quarterly and concrete enough to guide investment.

Metric Definition Why it matters
Decision coverage Weighted conservative coverage across registered decisions. Shows whether critical decisions have complete support.
False-confidence gap Weighted support score minus conservative coverage. Finds dashboards and alerts with missing ownership, action, or validation.
Decision latency Time from condition beginning to an accountable decision. Separates detection delay from interpretation, routing, approval, and action delay.
Customer-impact evidence latency Time until responders can state affected journeys, cohorts, and scope. Supports customer communication and prioritisation.
Alert precision and recall Fraction of alerts that map to significant events, and fraction of significant events that alert. Aligns with SLO alerting quality.
Evidence completeness Fraction of required fields present for incident, release, or journey decisions. Reveals missing correlation and instrumentation gaps.
Fallback availability Fraction of critical decisions with evidence access when primary dashboards fail. Protects incident response under dependency failure.
Feedback closure Fraction of postmortem and review actions completed or deliberately rejected by due date. Turns incidents into learning.
Decision outcome review Fraction of major decisions later reviewed for correctness and cost. Prevents confident but wrong routines from persisting.
Telemetry without decision Cost or volume of telemetry not mapped to any decision, investigation class, compliance need, or retention policy. Controls waste without cutting useful evidence blindly.

Decision latency deserves special attention. Mean time to recovery is useful but coarse. It often merges delays that require different fixes:

Decision latency = detection delay + routing delay + interpretation delay + authority delay + action delay + confirmation delay

A team with fast detection but slow authority does not need more alerts. It needs clearer permissions or a safer automated action. A team with slow interpretation may need better correlation, runbooks, or training. A team with slow confirmation may need better canary signals or customer-impact probes.

Practical measurement can start modestly. Select the ten most consequential recurring decisions for a service. Score coverage. Review the last three incidents or failed deployments. Record when useful evidence arrived, when the owner was identified, when mitigation was chosen, and when impact was confirmed. The purpose is to expose the weakest part of the decision loop, not to produce perfect accounting.

12. Where classic monitoring remains appropriate

The argument in this paper is not that classic monitoring is obsolete. It remains appropriate whenever the condition is known, the signal is reliable, and the action is clear.

Classic monitoring is especially useful for:

  • external availability checks for important endpoints;
  • SLO burn-rate alerts;
  • capacity and saturation thresholds;
  • queue age and backlog limits;
  • failed backups and restore tests;
  • certificate expiry;
  • batch completion and freshness;
  • data pipeline correctness probes;
  • security and access-control invariants;
  • budget, quota, and rate-limit exhaustion;
  • hardware, network, and platform health where failure modes are well understood.

The SRE book explicitly values black-box monitoring for paging because it forces attention to current user-visible symptoms, while white-box monitoring helps detect imminent problems and debug causes [4]. The right conclusion is to match the signal to the decision, rather than to choose black box over white box, or metrics over traces.

Classic monitoring also provides guardrails for automated systems. A release controller should not need an exploratory trace to know that a canary is burning error budget. An autoscaler should not wait for human analysis when saturation crosses a safe bound. A backup system should not need a dashboard review to report that a restore test failed.

The danger comes when classic monitoring is treated as complete observability. Known checks are necessary for known conditions. They do not eliminate the need for exploratory evidence, incident records, customer-impact analysis, or feedback loops.

13. Strongest counterargument

The strongest counterargument is that starting with decisions risks over-formalising a practice whose value often comes from exploration.

In a real outage, responders may not know the right decision in advance. They may begin with vague symptoms, discover unexpected dependencies, follow odd clues, and improvise around partial evidence. A rigid decision register could become another governance artifact that engineers ignore. Worse, it could narrow instrumentation to questions leaders already know how to ask, reducing the ability to discover unknown unknowns.

The counterargument also says telemetry foundations must come first. Without standardised collection, storage, trace propagation, and query capability, there is no evidence to route to any owner. Asking every team to map decisions before basic instrumentation exists may slow adoption. From this view, logs, metrics, traces, and dashboards are the platform. Decision quality emerges once the platform is available.

A decision-centred model can be misused. It can become compliance paperwork, privilege executive questions over engineering discovery, or make novel incidents look like exceptions rather than evidence that the model needs revision.

The answer is to treat decision coverage as a layer on top of telemetry foundations, not as a substitute. The platform still needs standard signals, context propagation, retention, sampling, query flexibility, and cost controls. The decision register should include exploratory decisions, such as: "Which hypothesis should the operations lead test next?", "Is the evidence sufficient to mitigate before root cause is known?", and "What data must be preserved before it expires?"

Unknown unknowns require flexible evidence. Decision coverage requires that flexible evidence reach someone who can use it. The two ideas are complementary when the model is used as a learning loop rather than a fixed checklist.

14. What this paper does not claim

This paper does not claim that telemetry is unimportant. Telemetry is the raw material of observability.

It does not claim that logs, metrics, traces, dashboards, profiles, or events are interchangeable. Each has strengths and limits.

It does not claim that every operational judgement can be specified in advance. Exploration remains central to debugging complex systems.

It does not claim that decision coverage is a scientifically validated industry benchmark. It is a proposed operating model that organisations should adapt and test.

It does not claim that buying or adopting OpenTelemetry creates observability by itself. OpenTelemetry improves collection, context, and interoperability; organisations still need ownership, interpretation, action, and feedback.

It does not claim that SLOs solve all reliability questions. Poor SLOs can miss user pain, hide cohorts, or become vanity metrics.

It does not claim that public incident reports provide complete evidence. Public reports are selective by necessity and should be read as useful evidence, not total truth.

It does not claim that alarm-fatigue findings from healthcare transfer directly to software operations. They are included only as supporting evidence that excessive, low-value alerts can weaken human response.

Conclusion

Observability should not begin with a list of tools. It should begin with the decisions that make the system safer, more reliable, more understandable, and more accountable.

Telemetry tells us what the system emitted. Observability asks whether the right person or controller can infer state, understand consequence, choose action, and learn from the result.

That shift matters for architecture. It changes what is instrumented, which attributes are standardised, how alerts are routed, how dashboards are designed, how SLOs are used, how incidents are managed, and how executives see reliability risk.

It also makes waste visible. A field that never supports a decision, investigation, compliance need, or learning loop may not deserve its cost. A dashboard that supports no owner may deserve retirement. An alert that produces no action may be harming the system it was meant to protect.

The useful question is not "Do we have logs, metrics, traces, and dashboards?"

It is "Which decisions are covered, which are not, and how do we know?"

About the author

Jason Doyle writes about reliable software, observability, applied AI, and practical controls for systems that influence human decisions. He publishes at jasondoyle.ie and can be contacted at contact@jasondoyle.ie.

References

  1. R. E. Kalman, On the General Theory of Control Systems, Proceedings of the First International Congress on Automatic Control, 1960.
  2. OpenTelemetry, Observability primer, https://opentelemetry.io/docs/concepts/observability-primer/.
  3. OpenTelemetry, Signals, https://opentelemetry.io/docs/concepts/signals/.
  4. Google SRE, Monitoring Distributed Systems, https://sre.google/sre-book/monitoring-distributed-systems/.
  5. Google SRE, Implementing SLOs, https://sre.google/workbook/implementing-slos/.
  6. Google SRE, Alerting on SLOs, https://sre.google/workbook/alerting-on-slos/.
  7. DORA, Monitoring and observability, https://dora.dev/capabilities/monitoring-and-observability/.
  8. DORA, DORA metrics, last updated 5 January 2026, https://dora.dev/guides/dora-metrics/.
  9. DORA, Accelerate State of DevOps Report 2024, last updated 13 April 2026, https://dora.dev/research/2024/dora-report/.
  10. DORA, User-centric focus, https://dora.dev/capabilities/user-centric-focus/.
  11. W3C, Trace Context Level 1, Recommendation, https://www.w3.org/TR/trace-context/.
  12. OpenTelemetry, Semantic Conventions, https://opentelemetry.io/docs/concepts/semantic-conventions/.
  13. Google Research, Dapper, a Large-Scale Distributed Systems Tracing Infrastructure, 2010, https://research.google/pubs/dapper-a-large-scale-distributed-systems-tracing-infrastructure/.
  14. Google SRE, Effective Troubleshooting, https://sre.google/sre-book/effective-troubleshooting/.
  15. Google SRE, Incident Response, https://sre.google/workbook/incident-response/.
  16. Google SRE, Postmortem Culture: Learning from Failure, https://sre.google/sre-book/postmortem-culture/.
  17. Slack Engineering, Slack's Outage on January 4th 2021, 2021, https://slack.engineering/slacks-outage-on-january-4th-2021/.
  18. Meta Engineering, More details about the October 4 outage, 2021, https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/.
  19. Cloudflare, Cloudflare outage on July 17, 2020, Cloudflare Blog, 2020, https://blog.cloudflare.com/cloudflare-outage-on-july-17-2020/.
  20. Atlassian, Post-Incident Review: April 2022 outage, Atlassian Blog, 2022, https://www.atlassian.com/blog/announcements/post-incident-review-april-2022-outage.
  21. GitLab, Postmortem of database outage of January 31, GitLab Blog, 2017, https://about.gitlab.com/blog/postmortem-of-database-outage-of-january-31/.
  22. NIST, Computer Security Incident Handling Guide, SP 800-61 Revision 2, 2012, https://doi.org/10.6028/NIST.SP.800-61r2.
  23. AHRQ, Making Healthcare Safer III: A Critical Analysis of Existing and Emerging Patient Safety Practices, alarm management safety culture, https://www.ncbi.nlm.nih.gov/books/NBK555522/.