Observability economics

The Observability Tax

When telemetry volume grows faster than decision value.

Jason Doyle First published: 1 October 2025 25 minute read

Disclosure: These views are my own and do not represent my current or any former employers.

Executive summary

This paper is the cost-and-stewardship companion to Observability Is a Decision System. That paper asks whether telemetry changes decisions. This one asks what those signals cost to keep, query, secure, explain, and eventually retire.

Observability is necessary. Modern software cannot be operated responsibly from hope, intuition, and scattered anecdotes. Teams need logs, metrics, traces, profiles, events, alerts, dashboards, and operational records. The problem is that telemetry has a growth model of its own: every service emits logs, every library adds metrics, every request can carry a trace, every incident adds an alert, and every unknown unknown becomes a reason to keep more data for longer.

The observability tax is the total cost imposed by telemetry that is collected, transported, stored, queried, maintained, reviewed, secured, and interpreted. It includes cloud bills and vendor invoices, but it is also network transfer, query compute, collector capacity, schema work, instrumentation labour, dashboard upkeep, alert triage, privacy review, and the attention consumed when humans try to make sense of too many signals.

The tax is worth paying when telemetry supports decisions. It is wasteful when telemetry survives because nobody knows who owns it, which decision it supports, what would happen if it were absent, or when it can safely expire.

OpenTelemetry makes standardised telemetry easier and more portable. Its purpose is to generate, collect, process, and export telemetry; it supports traces, metrics, logs, and baggage, with events and profiles also described in its signal documentation. It deliberately leaves storage and visualisation to other systems.[1][2] That is a strength. It also means the standard does not decide which evidence is worth keeping.

The public record supports the stewardship problem. CNCF material frames observability as a matter of purpose, instrumentation, and trade-off. OpenTelemetry sampling guidance recognises both cost reduction and the risk of missing critical evidence. Prometheus warns that label combinations create time series with RAM, CPU, disk, and network costs. Public cloud documentation shows billing categories beyond stored bytes. Industry surveys, while often vendor-produced, repeatedly report cost, complexity, tool sprawl, centralisation, and consolidation as live concerns. Google SRE and academic human-factors literature add the human constraint: pages, alarms, and information overload degrade judgement when signal is weak.[3][5][8][9][10][11][12][13][14][15][16][18][19][20]

The thesis is narrow: telemetry volume can grow faster than decision value. More data can make a system less understandable when ownership and decisions are unclear.

The proposal is to manage telemetry as a portfolio of evidence: define the decision, assign the owner, measure cost per supported decision, classify criticality, set budgets, preserve high-value evidence, and keep an explicit allocation for unknown-unknown discovery.

The goal is an observability system that makes the organisation more able to understand and act, without letting telemetry become an unmanaged standing tax.

1. Observability is not the same as telemetry volume

The companion paper argues that observability should be designed around the decisions it supports. This paper begins from the cost side of the same claim. If a signal has no owner, decision, retention class, or deletion rule, it is not free evidence. It is an obligation.

OpenTelemetry describes observability as understanding a system from its external outputs and focuses on generating, collecting, processing, and exporting telemetry data.[1] CNCF material makes the practical cloud native point: engineers must infer workload state from emitted signals.[3] Neither framing says that the most observable system is the one that emits the most data.

A system can produce terabytes of logs and remain obscure. It can expose a thousand dashboards while nobody knows which one should be trusted during an incident. It can trace every request and still fail to identify the owner of a degraded dependency. It can retain years of metrics and still lose the event that explains why a regulated decision was made.

Telemetry is evidence only when it can be used. The tax appears when teams add outputs faster than they add meaning, ownership, and deletion rules.

Most growth begins rationally. A team adds latency metrics. An incident exposes a logging gap. A migration creates a need for traces. A regulator asks for an audit trail. A platform emits its own events. Each step is defensible. The portfolio becomes expensive when nobody later asks whether the combined set still supports the decisions the organisation actually makes.

A single service may emit telemetry for incident response, product analysis, capacity planning, security detection, compliance evidence, support, billing, model monitoring, debugging, and executive reporting. Those uses have different users, tolerances, retention needs, privacy risks, and query patterns. Treating them as one undifferentiated stream turns a useful instrument into a standing expense.

2. What the tax includes

The visible cost of observability is usually the bill. That is only one layer.

Public pricing pages show the shape of the cost. Azure Monitor describes charges for log ingestion, retention, export, Prometheus samples ingested and query samples processed, alerts, notification choices, and time series created by query dimensions. Google Cloud Logging can charge for log storage and can charge again when the same entry is routed to multiple buckets. AWS CloudWatch pricing describes charges across ingestion, storage, monitored resources, diagnostic logs, metrics storage, and query use.[13][14][15]

The exact prices change. The pattern matters more than any current number: telemetry cost is not a single meter.

Storage

A log line is more than bytes on disk. It may be compressed, indexed, replicated, partitioned, encrypted, retained hot, copied cold, and routed to another system for security or audit. A metric is not one sample either: it is a sample for each time series, label combination, resolution, retention tier, and replica. A trace is a graph of spans whose value may depend on surrounding context.

Storage also creates future query work. Data that is cheap to retain but hard to find can still be expensive when needed.

Network, pipeline, and query capacity

Telemetry has to move. Agents, sidecars, libraries, collectors, queues, exporters, and backends consume CPU, memory, disk, and network capacity. OpenTelemetry recommends the Collector in many production scenarios because it can offload data quickly, batch, retry, encrypt, and filter sensitive data.[4] Those benefits come with components that must be deployed, scaled, monitored, secured, and upgraded.

Network costs are not limited to external egress. Internal fan-out, cross-region queries, duplicate export, and replay after failure also consume capacity. Uber's public M3 write-up describes downsampling at collection time partly to avoid the network, CPU, serialisation, and disk input/output tax of reading large volumes of stored metrics later.[25]

Querying is also a cost centre. A dashboard that refreshes every thirty seconds may issue hundreds of expensive queries per hour. During an incident, many engineers may run overlapping searches across high-cardinality data. Azure's documentation explicitly mentions Prometheus query samples processed and log-search alerts whose cost can depend on the number of time series created by dimensions.[13]

Engineering and cognitive cost

Telemetry must be designed. Prometheus says instrumentation should be an integral part of code and recommends metrics for libraries, online-serving systems, offline processing, batch jobs, failures, queues, caches, and collectors.[8] Google SRE material is explicit that monitoring a complex application is a significant engineering endeavour; even with substantial infrastructure, a Google SRE team of 10 to 12 members typically had one or two people primarily assigned to monitoring systems for the service.[16]

The final cost is attention. A page interrupts work, sleep, and personal time. Google SRE guidance warns that frequent pages cause people to second-guess, skim, or ignore alerts, sometimes masking real pages in noise.[16] Information-overload research reaches the same broad conclusion: people need goals, filtering, prioritisation, training, and organisational practices, not merely more information.[19]

3. Why volume grows faster than value

Telemetry grows because software grows. It also grows because of incentives. A developer who adds a log line pays almost none of its lifetime cost. A team that keeps an old dashboard avoids the immediate risk of deleting something somebody might need. A platform group may be rewarded for coverage rather than usefulness. A central budget may pay for telemetry that service teams emit.

The result is not bad faith so much as unmanaged common property.

High cardinality converts small choices into large systems

A metric with one label and ten values creates ten series. Add another label with ten values and the pessimistic upper bound becomes one hundred. Add user ID, request ID, build SHA, region, route, status, device, tenant, and feature flag, and a small line of code can create a large storage and query problem.

OpenTelemetry's .NET metrics guidance defines cardinality as the number of unique attribute combinations. It gives a simple example: seven attributes with thirty possible values each can lead to 21,870,000,000 combinations. It also identifies cardinality explosion as a known challenge that can create high costs or be used for denial of service, and notes a default cardinality limit of 2000 per metric in OpenTelemetry .NET.[7]

Prometheus guidance makes the operational warning direct: each unique labelset is an additional time series with RAM, CPU, disk, and network costs, and labels should not contain high-cardinality values such as user IDs, email addresses, or other unbounded sets.[9]

The question is not "is cardinality bad?" It is "which decisions require this cardinality, for how long, and under whose authority?"

Duplicate signals survive because deletion feels risky

Duplication often begins as migration safety. A team sends the same logs to an old platform and a new one, exports traces to two backends, or keeps Prometheus metrics while turning on OpenTelemetry. OpenTelemetry's Collector supports export to one or more backends and is intended to reduce the need to run and maintain multiple agents and collectors.[4] That flexibility is valuable. It can also make duplication easy unless export paths have owners and expiry dates.

The test is simple: if two systems store the same fact, what decision is better because both copies exist?

Legitimate answers include hot incident response plus cold audit, broad redacted access plus restricted evidence, high-cardinality debugging plus long-term aggregate trends, or temporary migration protection. "Nobody has asked us to turn it off" is not a legitimate architecture.

Dashboards become museums

Dashboards are cheap to create and expensive to trust. A useful dashboard answers a user's question. A dashboard museum preserves old questions with panels that still load and labels that still look official, even after the service, critical path, or alert route has changed.

Grafana Labs' survey material reports that complexity and overhead were the most cited concern in its fourth annual observability survey, and that 77 per cent of respondents said centralised observability saved time or money. A 2024 summary reported widespread use of multiple observability technologies and named cost as a major concern.[10][11]

These surveys are directional rather than neutral academic studies. Their value here is that the reported problems match the mechanics of the systems: more tools, data sources, dashboards, and overlapping queries make authority harder to identify.

4. Retention is an evidence decision

Retention policy often begins as a cost control. It should begin as an evidence decision.

The question goes beyond "how many days can we afford?" It includes which decisions the evidence supports, how soon those decisions are made, the consequence of absence, whether the data is personal or regulated, whether it can be aggregated or redacted, who can authorise deletion, and which legal, audit, incident, or customer-support hold can override expiry.

Uber's M3 case study is useful because it treats retention and resolution as explicit policies. M3 supported different retentions and granularities, such as short fine-grained retention and longer coarse-grained retention, selected by metric tag matching and storage policies.[25] The CNCF case study reports the same platform at large public scale: over 6.6 billion time series, 500 million metrics per second aggregated, 20 million metrics per second persisted, and storage that became 8.53 times more cost effective per metric per replica.[24]

The important lesson is not that other organisations should copy M3. Retention, rollup, and cost were treated as design problems, and that discipline is the transferable part.

Sampling is a trade-off, not a confession of failure

Sampling drops evidence. Keeping every successful low-latency request can also be wasteful. OpenTelemetry documentation takes both sides seriously: high-volume systems can often use sampling rates of one per cent or lower and still obtain a representative sample, but sampling carries compute cost, engineering cost, and opportunity cost if critical information is missed. Tail sampling can also be difficult to implement and operate because it may require stateful components with significant resources.[5]

Sampling should be governed by criticality rather than by volume alone. A trace for a routine health check can have different treatment from a trace that contains an error, high latency, payment failure, privilege change, regulated workflow, or customer-impacting exception. The decision to drop should itself be documented.

Deletion can be more dangerous than storage

Security incidents, customer disputes, regulated decisions, billing anomalies, fraud investigations, and serious outages may require evidence long after the first operational alert has closed. NIST SP 800-92 exists because log management is an enterprise security practice, not an incidental side effect of application development.[23]

This is no argument for retaining everything forever; deletion still needs to be managed as a decision with a failure mode.

A mature retention model keeps enough hot evidence to operate now, enough colder evidence to reconstruct important events later, and deletes or anonymises evidence that no longer has a legitimate purpose. The third point matters because GDPR Article 5 includes purpose limitation, data minimisation, storage limitation, integrity and confidentiality, and accountability.[22] OpenTelemetry's sensitive-data guidance says implementers are responsible for compliance, protecting sensitive information, obtaining necessary consents, and regularly reviewing attributes to ensure they remain necessary.[21]

Retention is therefore a balance between evidence loss and over-collection. Both can be governance failures.

5. The human tax

The observability tax is paid by machines first and people last. Machines can store more logs, scan more data, render more panels, and send more alerts. A human operator cannot expand attention in the same way.

Alert fatigue is a design problem

Google SRE guidance says alerts should not fire merely because something seems odd. Human alerts should be simple, reliable during incident response, and represent a clear failure.[16] The issue is design, not personal weakness: a system that repeatedly interrupts people with unactionable, duplicate, or low-context alerts trains them to distrust the signal.

Healthcare alarm-fatigue literature concerns a different domain, but it is useful because it studies repeated alarms in safety-critical settings. A 2025 scoping review found sustained work on definitions, influencing factors, consequences, and mitigation strategies.[18] Security operations research similarly treats SOC alert fatigue as a problem involving automation, augmentation, cognitive load, integration, and privacy compliance.[20]

A useful alert has an owner, a user-visible or business-visible symptom, a reason to interrupt now, enough context to begin action, a runbook, duplicate suppression, a review date, a measured false-alert rate, and incident history. An alert without those properties spends attention that may be needed later.

Instrumentation can become toil

SRE literature defines toil as work tied to running a production service that is manual, repetitive, automatable, tactical, without enduring value, and scaling linearly with service growth.[17]

Some observability work is engineering: a better semantic convention, reusable instrumentation library, removed class of duplicate alerts, or high-value signal on a critical path. Other work becomes toil: copying dashboard panels, editing thresholds nobody understands, adding logs to satisfy a template, responding to self-closing alerts, maintaining custom parsers, explaining monthly bill growth, or chasing owners for metrics emitted by retired services.

The tax can crowd out the work that would reduce it. A team overwhelmed by alert noise has less time to simplify alerting. A team fighting query cost has less time to improve schemas. A team instrumenting every new service by hand has less time to build safe defaults.

Dashboard sprawl creates false visibility

A dashboard is a promise: this view matters. When a dashboard is unowned, outdated, or overloaded, the promise becomes false. During an incident, the wrong dashboard can send engineers towards an old dependency, an obsolete threshold, or a metric whose labels changed during a migration.

Good dashboards should expire, just like data. Each should have a named audience, supported decision, owner, source queries, freshness expectations, known blind spots, review date, and links to runbooks and deeper evidence. A dashboard without a decision is documentation debt with live queries.

6. Ownership, privacy, and lock-in

Observability is organisational before it is technical. Telemetry crosses team boundaries: a trace may pass through several owners, a debugging log may contain a customer identifier needed by support and restricted by privacy policy, and a library metric label may multiply cost for every service that imports it. Without ownership, nobody can safely remove anything.

Data ownership is different from service ownership

The team that owns a service should not automatically own every decision made from that service's telemetry. A payment failure event, for example, may support incident response, finance reconciliation, customer support, fraud detection, regulatory reporting, product analysis, and vendor dispute resolution. Each use may require a different field, retention period, redaction method, access policy, and query interface.

A telemetry portfolio therefore needs both producers and decision owners. The producer knows how the data is generated. The decision owner knows why the data matters.

Privacy is not a post-processing feature

Telemetry often captures personal information accidentally: URLs contain email addresses, user IDs become metric attributes, headers contain tokens, error logs capture free text, traces record arguments, and profiles may expose file paths or data-dependent behaviour.

OpenTelemetry guidance says the implementer is responsible for sensitive data handling and should collect only data that serves an observability purpose, avoid personal information unless necessary, consider aggregated or anonymised data, and regularly review attributes. It also describes Collector processors for removing, filtering, redacting, hashing, or transforming attributes, while warning that hashing may not provide adequate anonymisation for small or predictable input spaces.[21]

It is better not to collect sensitive data than to collect it and hope a later processor removes it. Redaction is a control. It is not an excuse to ignore minimisation.

Open standards reduce lock-in, but do not remove switching cost

OpenTelemetry says one of its principles is that users own the data they generate and avoid vendor lock-in.[1] Grafana's survey reports that open source and open standards are important to 77 per cent of respondents, and that ease of switching vendors and backend technologies was one of the top concerns respondents hoped OpenTelemetry would help resolve.[10]

That progress does not eliminate switching cost. Lock-in can move from instrumentation to query languages, dashboard models, alert rules, retention policies, tail-sampling logic, derived fields, incident workflows, cost reports, access-control assumptions, and AI investigation features trained around one backend. OpenTelemetry sampling documentation notes that tail sampling often ends up as vendor-specific technology today when using paid vendors.[5]

That is a reason to keep the decision model portable, not a reason to avoid vendors altogether.

7. Measuring decision value

A telemetry strategy needs a unit of value. The proposed unit in this paper is the supported decision: a recurring or high-consequence choice that telemetry materially improves, such as paging, rollback, capacity increase, escalation, abuse blocking, SLO confirmation, retention hold, regulator notification, performance prioritisation, feature-flag removal, or false-positive rejection.

A signal that supports no decision may still have research value, but it should not hide inside the same budget as critical operational evidence.

Cost per supported decision

The simplest measurable concept is:

cost per supported decision =
  total telemetry cost for a portfolio / number of supported decisions

The numerator should include ingestion, storage, indexing, network transfer, query compute, collector and agent capacity, engineering maintenance, incident review time, alert handling, privacy work, duplicated export paths, and the opportunity cost of noisy signals. The denominator should not be dashboard views or query count. The useful question is whether the telemetry changed or justified an action.

Telemetry item Supported decision Review question
API latency histogram by route and status Page or rollback when users are affected Did it change response during recent incidents?
Per-user debug logs Investigate specific customer complaints Is access restricted and retention short?
Trace sample of successful requests Discover dependency drift Is the sample representative enough?
High-resolution CPU profiles Optimise a hotspot Is continuous collection needed or can it be triggered?
Long-term aggregate error rate Capacity and reliability planning Is rollup sufficient after 30 days?

A high cost per supported decision is not automatically bad. Some decisions are rare and critical. The point is to make the trade-off visible.

Evidence criticality

Telemetry should be classified by evidence criticality.

Class Description Default treatment
Critical evidence Needed to prove or reconstruct a regulated, security, safety, financial, or major customer-impacting event Preserve with access control, integrity controls, and documented retention
Operational evidence Needed for active incident response and short-term debugging Keep hot, searchable, and linked to runbooks
Diagnostic evidence Useful for rare defects, regressions, and performance analysis Sample, downsample, or move to cheaper storage after the active window
Trend evidence Useful for planning and historical comparison Aggregate and retain at lower resolution
Exploratory evidence Collected for discovery, hypothesis building, or unknown unknowns Time-box, sample, and require renewal
Redundant evidence Duplicated elsewhere without a distinct decision Remove or justify

This classification is more useful than signal type alone. A log can be critical evidence. A trace can be exploratory. A metric can be redundant. A profile can be operational during a performance incident and unnecessary after a fix.

Telemetry budgets

A telemetry budget is a limit tied to a decision portfolio, not merely a cloud spend target. It can cover daily ingest, active time series, maximum cardinality per metric, retained attributes, hot and cold retention, trace sampling by class, dashboard query cost, alert volume per on-call shift, false-alert rate, orphaned telemetry, and the percentage of signals with owners and review dates.

Budgets should be enforced near the source where possible. OpenTelemetry Views can customise which instruments are processed, which aggregation is used, and which attributes are reported. The metrics data model supports cost controls through temporal and spatial reaggregation. Collector processors can filter, transform, batch, and redact data before export.[4][6][7][21]

A budget that only appears on the invoice is too late.

Decision review

Review telemetry when the system, decision, cost, or risk changes. Ask which decisions it supported, which incidents used it, which alerts were actionable, which dashboards were used, which fields were never queried, which dimensions caused cardinality growth, which data was duplicated, which retention class is still justified, and which owners failed to respond.

If nobody can answer, the organisation does not have an observability system. It has a telemetry accumulation system.

8. Operating principles

These principles position the paper as a stewardship layer over the companion paper's decision-coverage model, rather than a second version of that model.

1. Start with costed decisions

Before adding or renewing a signal, write the decision it supports, the owner, the retention class, and the expected cost surface. The decision statement belongs in the companion model; this paper adds the budget, retention, privacy, and deletion questions.

2. Separate operational, analytical, and evidential uses

Incident response, product analytics, and compliance evidence may share a source event, but their derived forms can differ: hot operational data, aggregated trend data, restricted audit evidence, sampled diagnostic traces, or redacted support views. Each product needs its own owner.

3. Prefer bounded high value over broad low value

Google SRE's four golden signals - latency, traffic, errors, and saturation - are useful partly because they focus attention on symptoms users care about. [16] They establish a disciplined starting point, not a complete telemetry inventory. Add breadth when the decision value is clear.

4. Make cardinality a design review item

Metric labels, span attributes, log fields, and event dimensions should be reviewed before they enter production. The review should identify bounded and unbounded dimensions, expected values, maximum values, ownership, privacy classification, query-cost effect, the right signal type, and expiry for temporary dimensions. User ID may be a poor metric label and a necessary field in a restricted support log. The decision determines the treatment.

5. Treat dashboards and alerts as production assets

Dashboards and alerts should have ownership, review, source control where practical, validation for critical queries, change history, runbook links, retirement criteria, and incident feedback. A stale dashboard and a noisy alert both influence human decisions.

6. Keep an unknown-unknown allocation

Some telemetry should be collected because the organisation does not yet know what it will need. The mistake is letting exploratory telemetry become permanent without review. Give it a budget, sampling policy, privacy boundary, and expiry date. Renew it when it proves value. Retire it when it does not.

7. Preserve compact decision records

When telemetry supports a consequential decision, preserve the query or dashboard, time range, relevant evidence, decision maker, action, uncertainty, and later correction if any. The point is not to repeat the full decision-record model from the companion paper; it is to avoid keeping vast raw data while losing the smaller record that explains what the organisation believed at the time.

9. The strongest counterargument

The strongest counterargument is that the observability tax is the wrong frame.

Storage keeps getting cheaper. Compression improves. Cloud object stores, columnar formats, tiered retention, and open standards make it possible to keep large amounts of data at tolerable cost. Unknown unknowns are real. The most valuable evidence in an incident is often the field nobody predicted would matter. A strict decision-first model can prematurely remove weak signals that later explain a major failure.

This argument should be taken seriously.

Many severe incidents are understood only in retrospect. Early telemetry may look irrelevant because the organisation has not yet experienced the failure mode. Sampling can hide rare events. Aggregation can remove the outlier. Data minimisation can conflict with forensic reconstruction. A team that deletes unspecified evidence too aggressively may save money and lose the only path to truth.

There is also a political version of the argument. Requiring every signal to prove immediate value may favour established teams with known workflows and punish emerging services, research work, security hunting, and performance engineering. The absence of a current decision is not proof of uselessness.

The counterargument is strongest when it attacks false precision. Cost per supported decision can become theatre if teams invent decisions to defend their favourite data. Evidence criticality can be misclassified. Budgets can become blunt cuts. A cost-reduction programme can use the language of discipline while quietly deleting the evidence that would hold it accountable.

The answer is not to deny these risks.

The answer is to make unknown-unknown discovery explicit.

A mature telemetry portfolio should include exploratory and forensic classes. They should have budgets because they are real work, not because they are unimportant. They should have safeguards because they often contain sensitive or high-cardinality material. They should be reviewed because a permanent "just in case" stream is indistinguishable from an unowned stream.

Cheap storage changes the retention frontier. It does not remove network, index, query, privacy, ownership, or cognitive costs. It also does not guarantee that humans will find the evidence when needed.

The question is therefore not whether to keep only data with known immediate value. The question is whether every retained class of data has an explicit reason to exist.

Unknown unknowns are a reason.

They are not an exemption from design.

10. What this paper does not claim

This paper does not claim that observability is optional. Reliable systems need telemetry.

It does not claim that cost reduction should dominate incident response, security, safety, compliance, or customer trust.

It does not claim that high cardinality is always wrong. High-cardinality questions are often the questions that matter during debugging, abuse response, and customer support.

It does not claim that sampling is always safe. Sampling can remove the exact evidence needed later if the policy is poorly designed.

It does not claim that open source or OpenTelemetry automatically lowers cost. Open standards improve portability and ownership, but storage, query, retention, and operating costs remain.

It does not claim that vendor platforms are the problem. The tax exists in self-hosted systems too. Hardware, labour, complexity, and attention are still paid somewhere.

It does not claim that every signal must be tied to a known dashboard before it is collected. Exploration and forensic preservation are legitimate uses.

It does claim this:

Telemetry should be governed by decision value, evidence criticality, and human understandability. More data is not the same as more observability.

Conclusion

Observability began as a way to make complex systems understandable. It can fail by becoming another source of complexity.

The companion paper asks whether evidence reaches a decision. This paper asks whether the evidence portfolio is worth its storage, query, engineering, privacy, and attention costs.

The failure usually arrives as small additions: a field on every log, a label on every metric, a trace kept for every request, a copied dashboard, a duplicate pipeline after migration, an alert added after an incident and never reviewed. Each addition is defensible. The portfolio becomes expensive.

What the observability tax needs is stewardship, not austerity.

Treat telemetry as evidence. Ask what decision it supports. Classify its criticality. Budget its volume, cardinality, query load, retention, and human attention. Preserve what must be preserved. Sample or aggregate what can be sampled or aggregated. Delete what has no purpose. Keep exploratory capacity, but give it a name and a review date.

An observable system is not the system that remembers everything. It is the system that keeps enough of the right evidence, in a usable form, for the people and machines responsible for making decisions.

About the author

Jason Doyle writes about reliable software, observability, applied AI, and practical controls for systems that influence human decisions. He publishes at jasondoyle.ie and can be contacted at contact@jasondoyle.ie.

References

  1. OpenTelemetry, What is OpenTelemetry?, https://opentelemetry.io/docs/what-is-opentelemetry/.
  2. OpenTelemetry, Signals, https://opentelemetry.io/docs/concepts/signals/.
  3. CNCF TAG Observability, Observability Whitepaper, 2023, https://github.com/cncf/tag-observability/blob/main/whitepaper.md.
  4. OpenTelemetry, Collector, https://opentelemetry.io/docs/collector/.
  5. OpenTelemetry, Sampling, https://opentelemetry.io/docs/concepts/sampling/.
  6. OpenTelemetry, Metrics Data Model, https://opentelemetry.io/docs/specs/otel/metrics/data-model/.
  7. OpenTelemetry, .NET metrics best practices, https://opentelemetry.io/docs/languages/dotnet/metrics/best-practices/.
  8. Prometheus, Instrumentation, https://prometheus.io/docs/practices/instrumentation/.
  9. Prometheus, Metric and label naming, https://prometheus.io/docs/practices/naming/.
  10. Grafana Labs, 4th annual Observability Survey, https://grafana.com/observability-survey/.
  11. Grafana Labs, 5 key takeaways from the Grafana Labs' 2024 Observability Survey, 2024, https://grafana.com/blog/5-key-takeaways-from-the-grafana-labs-2024-observability-survey/.
  12. New Relic, 2025 Observability Forecast press release, 17 September 2025, https://newrelic.com/press-release/20250917.
  13. Microsoft Learn, Azure Monitor cost and usage, archived source revision, 29 July 2025, https://github.com/MicrosoftDocs/azure-monitor-docs/blob/5b79f1a71a2523e964d70bbd7e9980e61d7d453a/articles/azure-monitor/fundamentals/cost-usage.md.
  14. Google Cloud, Google Cloud Observability pricing, https://cloud.google.com/products/observability/pricing.
  15. AWS, Amazon CloudWatch pricing, https://aws.amazon.com/cloudwatch/pricing/.
  16. Rob Ewaschuk, Monitoring Distributed Systems, Google SRE, https://sre.google/sre-book/monitoring-distributed-systems/.
  17. Vivek Rau, Eliminating Toil, Google SRE, https://sre.google/sre-book/eliminating-toil/.
  18. Elizabeth Anna Mathilde Michels et al., Alarm fatigue in healthcare, BMC Nursing, 2025, https://doi.org/10.1186/s12912-025-03369-2.
  19. Miriam Arnold, Mascha Goldschmitt, and Thomas Rigotti, Dealing with information overload, Frontiers in Psychology, 2023, https://doi.org/10.3389/fpsyg.2023.1122200.
  20. ACM Digital Library, Alert Fatigue in Security Operations Centres, 2025, https://doi.org/10.1145/3723158.
  21. OpenTelemetry, Handling sensitive data, https://opentelemetry.io/docs/security/handling-sensitive-data/.
  22. European Union, Article 5 GDPR, https://gdpr-info.eu/art-5-gdpr/.
  23. NIST, SP 800-92: Guide to Computer Security Log Management, 2006, https://csrc.nist.gov/pubs/sp/800/92/final.
  24. CNCF, How Uber is monitoring 4,000 microservices, https://www.cncf.io/case-studies/uber/.
  25. Uber Engineering, M3: Uber's Open Source, Large-scale Metrics Platform for Prometheus, 2018, https://www.uber.com/us/en/blog/m3/.