> Disclosure: These views are my own and do not represent my current or any former employers.

# The Observability Tax

## When telemetry volume grows faster than decision value

Author: Jason Doyle

First published: 1 October 2025

Source review updated: 25 August 2026

## Executive summary

This paper is the cost-and-stewardship companion to
[Observability Is a Decision System](/whitepapers/observability-is-a-decision-system/).
That paper asks whether telemetry changes decisions. This one asks what those
signals cost to keep, query, secure, explain, and eventually retire.

Observability is necessary. Modern software cannot be operated responsibly from
hope, intuition, and scattered anecdotes. Teams need logs, metrics, traces,
profiles, events, alerts, dashboards, and operational records. The problem is
that telemetry has a growth model of its own: every service emits logs, every
library adds metrics, every request can carry a trace, every incident adds an
alert, and every unknown unknown becomes a reason to keep more data for longer.

The observability tax is the total cost imposed by telemetry that is collected,
transported, stored, queried, maintained, reviewed, secured, and interpreted.
It includes cloud bills and vendor invoices, but it is also network transfer,
query compute, collector capacity, schema work, instrumentation labour,
dashboard upkeep, alert triage, privacy review, and the attention consumed when
humans try to make sense of too many signals.

The tax is worth paying when telemetry supports decisions. It is wasteful when
telemetry survives because nobody knows who owns it, which decision it supports,
what would happen if it were absent, or when it can safely expire.

OpenTelemetry makes standardised telemetry easier and more portable. Its
purpose is to generate, collect, process, and export telemetry; it supports
traces, metrics, logs, and baggage, with events and profiles also described in
its signal documentation. It deliberately leaves storage and visualisation to
other systems.\[1\]\[2\] That is a strength. It also means the standard does
not decide which evidence is worth keeping.

The public record supports the stewardship problem. CNCF material frames
observability as a matter of purpose, instrumentation, and trade-off.
OpenTelemetry sampling guidance recognises both cost reduction and the risk of
missing critical evidence. Prometheus warns that label combinations create time
series with RAM, CPU, disk, and network costs. Public cloud documentation shows
billing categories beyond stored bytes. Industry surveys, while often
vendor-produced, repeatedly report cost, complexity, tool sprawl,
centralisation, and consolidation as live concerns. Google SRE and academic
human-factors literature add the human constraint: pages, alarms, and
information overload degrade judgement when signal is weak.\[3\]\[5\]\[8\]\[9\]\[10\]\[11\]\[12\]\[13\]\[14\]\[15\]\[16\]\[18\]\[19\]\[20\]

The thesis is narrow: telemetry volume can grow faster than decision value.
More data can make a system less understandable when ownership and decisions are
unclear.

The proposal is to manage telemetry as a portfolio of evidence: define the
decision, assign the owner, measure cost per supported decision, classify
criticality, set budgets, preserve high-value evidence, and keep an explicit
allocation for unknown-unknown discovery.

The goal is an observability system that makes the organisation more able to
understand and act, without letting telemetry become an unmanaged standing tax.

## 1. Observability is not the same as telemetry volume

The companion paper argues that observability should be designed around the
decisions it supports. This paper begins from the cost side of the same claim.
If a signal has no owner, decision, retention class, or deletion rule, it is not
free evidence. It is an obligation.

OpenTelemetry describes observability as understanding a system from its
external outputs and focuses on generating, collecting, processing, and
exporting telemetry data.\[1\] CNCF material makes the practical cloud native
point: engineers must infer workload state from emitted signals.\[3\] Neither
framing says that the most observable system is the one that emits the most
data.

A system can produce terabytes of logs and remain obscure. It can expose a
thousand dashboards while nobody knows which one should be trusted during an
incident. It can trace every request and still fail to identify the owner of a
degraded dependency. It can retain years of metrics and still lose the event
that explains why a regulated decision was made.

Telemetry is evidence only when it can be used. The tax appears when teams add
outputs faster than they add meaning, ownership, and deletion rules.

Most growth begins rationally. A team adds latency metrics. An incident exposes
a logging gap. A migration creates a need for traces. A regulator asks for an
audit trail. A platform emits its own events. Each step is defensible. The
portfolio becomes expensive when nobody later asks whether the combined set
still supports the decisions the organisation actually makes.

A single service may emit telemetry for incident response, product analysis,
capacity planning, security detection, compliance evidence, support, billing,
model monitoring, debugging, and executive reporting. Those uses have different
users, tolerances, retention needs, privacy risks, and query patterns. Treating
them as one undifferentiated stream turns a useful instrument into a standing
expense.

## 2. What the tax includes

The visible cost of observability is usually the bill. That is only one layer.

Public pricing pages show the shape of the cost. Azure Monitor describes
charges for log ingestion, retention, export, Prometheus samples ingested and
query samples processed, alerts, notification choices, and time series created
by query dimensions. Google Cloud Logging can charge for log storage and can
charge again when the same entry is routed to multiple buckets. AWS CloudWatch
pricing describes charges across ingestion, storage, monitored resources,
diagnostic logs, metrics storage, and query use.\[13\]\[14\]\[15\]

The exact prices change. The pattern matters more than any current number:
telemetry cost is not a single meter.

### Storage

A log line is more than bytes on disk. It may be compressed, indexed, replicated,
partitioned, encrypted, retained hot, copied cold, and routed to another system
for security or audit. A metric is not one sample either: it is a sample for each time
series, label combination, resolution, retention tier, and replica. A trace is a
graph of spans whose value may depend on surrounding context.

Storage also creates future query work. Data that is cheap to retain but hard
to find can still be expensive when needed.

### Network, pipeline, and query capacity

Telemetry has to move. Agents, sidecars, libraries, collectors, queues,
exporters, and backends consume CPU, memory, disk, and network capacity.
OpenTelemetry recommends the Collector in many production scenarios because it
can offload data quickly, batch, retry, encrypt, and filter sensitive data.\[4\]
Those benefits come with components that must be deployed, scaled, monitored,
secured, and upgraded.

Network costs are not limited to external egress. Internal fan-out, cross-region
queries, duplicate export, and replay after failure also consume capacity.
Uber's public M3 write-up describes downsampling at collection time partly to
avoid the network, CPU, serialisation, and disk input/output tax of reading
large volumes of stored metrics later.\[25\]

Querying is also a cost centre. A dashboard that refreshes every thirty seconds
may issue hundreds of expensive queries per hour. During an incident, many
engineers may run overlapping searches across high-cardinality data. Azure's
documentation explicitly mentions Prometheus query samples processed and
log-search alerts whose cost can depend on the number of time series created by
dimensions.\[13\]

### Engineering and cognitive cost

Telemetry must be designed. Prometheus says instrumentation should be an
integral part of code and recommends metrics for libraries, online-serving
systems, offline processing, batch jobs, failures, queues, caches, and
collectors.\[8\] Google SRE material is explicit that monitoring a complex
application is a significant engineering endeavour; even with substantial
infrastructure, a Google SRE team of 10 to 12 members typically had one or two
people primarily assigned to monitoring systems for the service.\[16\]

The final cost is attention. A page interrupts work, sleep, and personal time.
Google SRE guidance warns that frequent pages cause people to second-guess,
skim, or ignore alerts, sometimes masking real pages in noise.\[16\]
Information-overload research reaches the same broad conclusion: people need
goals, filtering, prioritisation, training, and organisational practices, not
merely more information.\[19\]

## 3. Why volume grows faster than value

Telemetry grows because software grows. It also grows because of incentives.
A developer who adds a log line pays almost none of its lifetime cost. A team
that keeps an old dashboard avoids the immediate risk of deleting something
somebody might need. A platform group may be rewarded for coverage rather than
usefulness. A central budget may pay for telemetry that service teams emit.

The result is not bad faith so much as unmanaged common property.

### High cardinality converts small choices into large systems

A metric with one label and ten values creates ten series. Add another label
with ten values and the pessimistic upper bound becomes one hundred. Add user
ID, request ID, build SHA, region, route, status, device, tenant, and feature
flag, and a small line of code can create a large storage and query problem.

OpenTelemetry's .NET metrics guidance defines cardinality as the number of
unique attribute combinations. It gives a simple example: seven attributes with
thirty possible values each can lead to 21,870,000,000 combinations. It also
identifies cardinality explosion as a known challenge that can create high costs
or be used for denial of service, and notes a default cardinality limit of 2000
per metric in OpenTelemetry .NET.\[7\]

Prometheus guidance makes the operational warning direct: each unique labelset
is an additional time series with RAM, CPU, disk, and network costs, and labels
should not contain high-cardinality values such as user IDs, email addresses, or
other unbounded sets.\[9\]

The question is not "is cardinality bad?" It is "which decisions require this
cardinality, for how long, and under whose authority?"

### Duplicate signals survive because deletion feels risky

Duplication often begins as migration safety. A team sends the same logs to an
old platform and a new one, exports traces to two backends, or keeps Prometheus
metrics while turning on OpenTelemetry. OpenTelemetry's Collector supports
export to one or more backends and is intended to reduce the need to run and
maintain multiple agents and collectors.\[4\] That flexibility is valuable. It
can also make duplication easy unless export paths have owners and expiry dates.

The test is simple: if two systems store the same fact, what decision is better
because both copies exist?

Legitimate answers include hot incident response plus cold audit, broad redacted
access plus restricted evidence, high-cardinality debugging plus long-term
aggregate trends, or temporary migration protection. "Nobody has asked us to
turn it off" is not a legitimate architecture.

### Dashboards become museums

Dashboards are cheap to create and expensive to trust. A useful dashboard
answers a user's question. A dashboard museum preserves old questions with
panels that still load and labels that still look official, even after the
service, critical path, or alert route has changed.

Grafana Labs' survey material reports that complexity and overhead were the most
cited concern in its fourth annual observability survey, and that 77 per cent of
respondents said centralised observability saved time or money. A 2024 summary
reported widespread use of multiple observability technologies and named cost as
a major concern.\[10\]\[11\]

These surveys are directional rather than neutral academic studies. Their value
here is that the reported problems match the mechanics of the systems: more
tools, data sources, dashboards, and overlapping queries make authority harder
to identify.

## 4. Retention is an evidence decision

Retention policy often begins as a cost control. It should begin as an evidence
decision.

The question goes beyond "how many days can we afford?" It includes which
decisions the evidence supports, how soon those decisions are made, the
consequence of absence, whether the data is personal or regulated, whether it
can be aggregated or redacted, who can authorise deletion, and which legal,
audit, incident, or customer-support hold can override expiry.

Uber's M3 case study is useful because it treats retention and resolution as
explicit policies. M3 supported different retentions and granularities, such as
short fine-grained retention and longer coarse-grained retention, selected by
metric tag matching and storage policies.\[25\] The CNCF case study reports the same
platform at large public scale: over 6.6 billion time series, 500 million
metrics per second aggregated, 20 million metrics per second persisted, and
storage that became 8.53 times more cost effective per metric per replica.\[24\]

The important lesson is not that other organisations should copy M3. Retention,
rollup, and cost were treated as design problems, and that discipline is the
transferable part.

### Sampling is a trade-off, not a confession of failure

Sampling drops evidence. Keeping every successful low-latency request can also
be wasteful. OpenTelemetry documentation takes both sides seriously: high-volume
systems can often use sampling rates of one per cent or lower and still obtain a
representative sample, but sampling carries compute cost, engineering cost, and
opportunity cost if critical information is missed. Tail sampling can also be
difficult to implement and operate because it may require stateful components
with significant resources.\[5\]

Sampling should be governed by criticality rather than by volume alone. A trace
for a routine health check can have different treatment from a trace that
contains an error, high latency, payment failure, privilege change, regulated
workflow, or customer-impacting exception. The decision to drop should itself
be documented.

### Deletion can be more dangerous than storage

Security incidents, customer disputes, regulated decisions, billing anomalies,
fraud investigations, and serious outages may require evidence long after the
first operational alert has closed. NIST SP 800-92 exists because log
management is an enterprise security practice, not an incidental side effect of
application development.\[23\]

This is no argument for retaining everything forever; deletion still needs to be
managed as a decision with a failure mode.

A mature retention model keeps enough hot evidence to operate now, enough colder
evidence to reconstruct important events later, and deletes or anonymises
evidence that no longer has a legitimate purpose. The third point matters
because GDPR Article 5 includes purpose limitation, data minimisation, storage
limitation, integrity and confidentiality, and accountability.\[22\]
OpenTelemetry's sensitive-data guidance says implementers are responsible for
compliance, protecting sensitive information, obtaining necessary consents, and
regularly reviewing attributes to ensure they remain necessary.\[21\]

Retention is therefore a balance between evidence loss and over-collection.
Both can be governance failures.

## 5. The human tax

The observability tax is paid by machines first and people last. Machines can
store more logs, scan more data, render more panels, and send more alerts. A
human operator cannot expand attention in the same way.

### Alert fatigue is a design problem

Google SRE guidance says alerts should not fire merely because something seems
odd. Human alerts should be simple, reliable during incident response, and
represent a clear failure.\[16\] The issue is design, not personal weakness: a
system that repeatedly interrupts people with unactionable, duplicate, or
low-context alerts trains them to distrust the signal.

Healthcare alarm-fatigue literature concerns a different domain, but it is
useful because it studies repeated alarms in safety-critical settings. A 2025
scoping review found sustained work on definitions, influencing factors,
consequences, and mitigation strategies.\[18\] Security operations research
similarly treats SOC alert fatigue as a problem involving automation,
augmentation, cognitive load, integration, and privacy compliance.\[20\]

A useful alert has an owner, a user-visible or business-visible symptom, a
reason to interrupt now, enough context to begin action, a runbook, duplicate
suppression, a review date, a measured false-alert rate, and incident history.
An alert without those properties spends attention that may be needed later.

### Instrumentation can become toil

SRE literature defines toil as work tied to running a production service that is
manual, repetitive, automatable, tactical, without enduring value, and scaling
linearly with service growth.\[17\]

Some observability work is engineering: a better semantic convention, reusable
instrumentation library, removed class of duplicate alerts, or high-value signal
on a critical path. Other work becomes toil: copying dashboard panels, editing
thresholds nobody understands, adding logs to satisfy a template, responding to
self-closing alerts, maintaining custom parsers, explaining monthly bill growth,
or chasing owners for metrics emitted by retired services.

The tax can crowd out the work that would reduce it. A team overwhelmed by
alert noise has less time to simplify alerting. A team fighting query cost has
less time to improve schemas. A team instrumenting every new service by hand has
less time to build safe defaults.

### Dashboard sprawl creates false visibility

A dashboard is a promise: this view matters. When a dashboard is unowned,
outdated, or overloaded, the promise becomes false. During an incident, the
wrong dashboard can send engineers towards an old dependency, an obsolete
threshold, or a metric whose labels changed during a migration.

Good dashboards should expire, just like data. Each should have a named
audience, supported decision, owner, source queries, freshness expectations,
known blind spots, review date, and links to runbooks and deeper evidence. A
dashboard without a decision is documentation debt with live queries.

## 6. Ownership, privacy, and lock-in

Observability is organisational before it is technical. Telemetry crosses team
boundaries: a trace may pass through several owners, a debugging log may contain
a customer identifier needed by support and restricted by privacy policy, and a
library metric label may multiply cost for every service that imports it.
Without ownership, nobody can safely remove anything.

### Data ownership is different from service ownership

The team that owns a service should not automatically own every decision made
from that service's telemetry. A payment failure event, for example, may support
incident response, finance reconciliation, customer support, fraud detection,
regulatory reporting, product analysis, and vendor dispute resolution. Each use
may require a different field, retention period, redaction method, access
policy, and query interface.

A telemetry portfolio therefore needs both producers and decision owners. The
producer knows how the data is generated. The decision owner knows why the data
matters.

### Privacy is not a post-processing feature

Telemetry often captures personal information accidentally: URLs contain email
addresses, user IDs become metric attributes, headers contain tokens, error logs
capture free text, traces record arguments, and profiles may expose file paths
or data-dependent behaviour.

OpenTelemetry guidance says the implementer is responsible for sensitive data
handling and should collect only data that serves an observability purpose,
avoid personal information unless necessary, consider aggregated or anonymised
data, and regularly review attributes. It also describes Collector processors
for removing, filtering, redacting, hashing, or transforming attributes, while
warning that hashing may not provide adequate anonymisation for small or
predictable input spaces.\[21\]

It is better not to collect sensitive data than to collect it and hope a later
processor removes it. Redaction is a control. It is not an excuse to ignore
minimisation.

### Open standards reduce lock-in, but do not remove switching cost

OpenTelemetry says one of its principles is that users own the data they
generate and avoid vendor lock-in.\[1\] Grafana's survey reports that open
source and open standards are important to 77 per cent of respondents, and that
ease of switching vendors and backend technologies was one of the top concerns
respondents hoped OpenTelemetry would help resolve.\[10\]

That progress does not eliminate switching cost. Lock-in can move from
instrumentation to query languages, dashboard models, alert rules, retention
policies, tail-sampling logic, derived fields, incident workflows, cost reports,
access-control assumptions, and AI investigation features trained around one
backend. OpenTelemetry sampling documentation notes that tail sampling often
ends up as vendor-specific technology today when using paid vendors.\[5\]

That is a reason to keep the decision model portable, not a reason to avoid
vendors altogether.

## 7. Measuring decision value

A telemetry strategy needs a unit of value. The proposed unit in this paper is
the supported decision: a recurring or high-consequence choice that telemetry
materially improves, such as paging, rollback, capacity increase, escalation,
abuse blocking, SLO confirmation, retention hold, regulator notification,
performance prioritisation, feature-flag removal, or false-positive rejection.

A signal that supports no decision may still have research value, but it should
not hide inside the same budget as critical operational evidence.

### Cost per supported decision

The simplest measurable concept is:

```text
cost per supported decision =
  total telemetry cost for a portfolio / number of supported decisions
```

The numerator should include ingestion, storage, indexing, network transfer,
query compute, collector and agent capacity, engineering maintenance, incident
review time, alert handling, privacy work, duplicated export paths, and the
opportunity cost of noisy signals. The denominator should not be dashboard views
or query count. The useful question is whether the telemetry changed or
justified an action.

| Telemetry item | Supported decision | Review question |
| --- | --- | --- |
| API latency histogram by route and status | Page or rollback when users are affected | Did it change response during recent incidents? |
| Per-user debug logs | Investigate specific customer complaints | Is access restricted and retention short? |
| Trace sample of successful requests | Discover dependency drift | Is the sample representative enough? |
| High-resolution CPU profiles | Optimise a hotspot | Is continuous collection needed or can it be triggered? |
| Long-term aggregate error rate | Capacity and reliability planning | Is rollup sufficient after 30 days? |

A high cost per supported decision is not automatically bad. Some decisions are
rare and critical. The point is to make the trade-off visible.

### Evidence criticality

Telemetry should be classified by evidence criticality.

| Class | Description | Default treatment |
| --- | --- | --- |
| Critical evidence | Needed to prove or reconstruct a regulated, security, safety, financial, or major customer-impacting event | Preserve with access control, integrity controls, and documented retention |
| Operational evidence | Needed for active incident response and short-term debugging | Keep hot, searchable, and linked to runbooks |
| Diagnostic evidence | Useful for rare defects, regressions, and performance analysis | Sample, downsample, or move to cheaper storage after the active window |
| Trend evidence | Useful for planning and historical comparison | Aggregate and retain at lower resolution |
| Exploratory evidence | Collected for discovery, hypothesis building, or unknown unknowns | Time-box, sample, and require renewal |
| Redundant evidence | Duplicated elsewhere without a distinct decision | Remove or justify |

This classification is more useful than signal type alone. A log can be
critical evidence. A trace can be exploratory. A metric can be redundant. A
profile can be operational during a performance incident and unnecessary after a
fix.

### Telemetry budgets

A telemetry budget is a limit tied to a decision portfolio, not merely a cloud
spend target. It can cover daily ingest, active time series, maximum cardinality
per metric, retained attributes, hot and cold retention, trace sampling by
class, dashboard query cost, alert volume per on-call shift, false-alert rate,
orphaned telemetry, and the percentage of signals with owners and review dates.

Budgets should be enforced near the source where possible. OpenTelemetry Views
can customise which instruments are processed, which aggregation is used, and
which attributes are reported. The metrics data model supports cost controls
through temporal and spatial reaggregation. Collector processors can filter,
transform, batch, and redact data before export.\[4\]\[6\]\[7\]\[21\]

A budget that only appears on the invoice is too late.

### Decision review

Review telemetry when the system, decision, cost, or risk changes. Ask which
decisions it supported, which incidents used it, which alerts were actionable,
which dashboards were used, which fields were never queried, which dimensions
caused cardinality growth, which data was duplicated, which retention class is
still justified, and which owners failed to respond.

If nobody can answer, the organisation does not have an observability system.
It has a telemetry accumulation system.

## 8. Operating principles

These principles position the paper as a stewardship layer over the companion
paper's decision-coverage model, rather than a second version of that model.

### 1. Start with costed decisions

Before adding or renewing a signal, write the decision it supports, the owner,
the retention class, and the expected cost surface. The decision statement
belongs in the companion model; this paper adds the budget, retention, privacy,
and deletion questions.

### 2. Separate operational, analytical, and evidential uses

Incident response, product analytics, and compliance evidence may share a source
event, but their derived forms can differ: hot operational data, aggregated
trend data, restricted audit evidence, sampled diagnostic traces, or redacted
support views. Each product needs its own owner.

### 3. Prefer bounded high value over broad low value

Google SRE's four golden signals - latency, traffic, errors, and saturation -
are useful partly because they focus attention on symptoms users care about.
\[16\] They establish a disciplined starting point, not a complete telemetry
inventory. Add breadth when the decision value is clear.

### 4. Make cardinality a design review item

Metric labels, span attributes, log fields, and event dimensions should be
reviewed before they enter production. The review should identify bounded and
unbounded dimensions, expected values, maximum values, ownership, privacy
classification, query-cost effect, the right signal type, and expiry for
temporary dimensions. User ID may be a poor metric label and a necessary field
in a restricted support log. The decision determines the treatment.

### 5. Treat dashboards and alerts as production assets

Dashboards and alerts should have ownership, review, source control where
practical, validation for critical queries, change history, runbook links,
retirement criteria, and incident feedback. A stale dashboard and a noisy alert
both influence human decisions.

### 6. Keep an unknown-unknown allocation

Some telemetry should be collected because the organisation does not yet know
what it will need. The mistake is letting exploratory telemetry become
permanent without review. Give it a budget, sampling policy, privacy boundary,
and expiry date. Renew it when it proves value. Retire it when it does not.

### 7. Preserve compact decision records

When telemetry supports a consequential decision, preserve the query or
dashboard, time range, relevant evidence, decision maker, action, uncertainty,
and later correction if any. The point is not to repeat the full decision-record
model from the companion paper; it is to avoid keeping vast raw data while
losing the smaller record that explains what the organisation believed at the
time.

## 9. The strongest counterargument

The strongest counterargument is that the observability tax is the wrong frame.

Storage keeps getting cheaper. Compression improves. Cloud object stores,
columnar formats, tiered retention, and open standards make it possible to keep
large amounts of data at tolerable cost. Unknown unknowns are real. The most
valuable evidence in an incident is often the field nobody predicted would
matter. A strict decision-first model can prematurely remove weak signals that
later explain a major failure.

This argument should be taken seriously.

Many severe incidents are understood only in retrospect. Early telemetry may
look irrelevant because the organisation has not yet experienced the failure
mode. Sampling can hide rare events. Aggregation can remove the outlier. Data
minimisation can conflict with forensic reconstruction. A team that deletes
unspecified evidence too aggressively may save money and lose the only path to
truth.

There is also a political version of the argument. Requiring every signal to
prove immediate value may favour established teams with known workflows and
punish emerging services, research work, security hunting, and performance
engineering. The absence of a current decision is not proof of uselessness.

The counterargument is strongest when it attacks false precision. Cost per
supported decision can become theatre if teams invent decisions to defend their
favourite data. Evidence criticality can be misclassified. Budgets can become
blunt cuts. A cost-reduction programme can use the language of discipline while
quietly deleting the evidence that would hold it accountable.

The answer is not to deny these risks.

The answer is to make unknown-unknown discovery explicit.

A mature telemetry portfolio should include exploratory and forensic classes.
They should have budgets because they are real work, not because they are
unimportant. They should have safeguards because they often contain sensitive or
high-cardinality material. They should be reviewed because a permanent
"just in case" stream is indistinguishable from an unowned stream.

Cheap storage changes the retention frontier. It does not remove network,
index, query, privacy, ownership, or cognitive costs. It also does not guarantee
that humans will find the evidence when needed.

The question is therefore not whether to keep only data with known immediate
value. The question is whether every retained class of data has an explicit
reason to exist.

Unknown unknowns are a reason.

They are not an exemption from design.

## 10. What this paper does not claim

This paper does not claim that observability is optional. Reliable systems need
telemetry.

It does not claim that cost reduction should dominate incident response,
security, safety, compliance, or customer trust.

It does not claim that high cardinality is always wrong. High-cardinality
questions are often the questions that matter during debugging, abuse response,
and customer support.

It does not claim that sampling is always safe. Sampling can remove the exact
evidence needed later if the policy is poorly designed.

It does not claim that open source or OpenTelemetry automatically lowers cost.
Open standards improve portability and ownership, but storage, query,
retention, and operating costs remain.

It does not claim that vendor platforms are the problem. The tax exists in
self-hosted systems too. Hardware, labour, complexity, and attention are still
paid somewhere.

It does not claim that every signal must be tied to a known dashboard before it
is collected. Exploration and forensic preservation are legitimate uses.

It does claim this:

Telemetry should be governed by decision value, evidence criticality, and human
understandability. More data is not the same as more observability.

## Conclusion

Observability began as a way to make complex systems understandable. It can
fail by becoming another source of complexity.

The companion paper asks whether evidence reaches a decision. This paper asks
whether the evidence portfolio is worth its storage, query, engineering,
privacy, and attention costs.

The failure usually arrives as small additions: a field on every log, a label on
every metric, a trace kept for every request, a copied dashboard, a duplicate
pipeline after migration, an alert added after an incident and never reviewed.
Each addition is defensible. The portfolio becomes expensive.

What the observability tax needs is stewardship, not austerity.

Treat telemetry as evidence. Ask what decision it supports. Classify its
criticality. Budget its volume, cardinality, query load, retention, and human
attention. Preserve what must be preserved. Sample or aggregate what can be
sampled or aggregated. Delete what has no purpose. Keep exploratory capacity,
but give it a name and a review date.

An observable system is not the system that remembers everything. It is the
system that keeps enough of the right evidence, in a usable form, for the people
and machines responsible for making decisions.



## About the author

Jason Doyle writes about reliable software, observability, applied AI, and
practical controls for systems that influence human decisions. He publishes at
[jasondoyle.ie](https://jasondoyle.ie) and can be contacted at
[contact@jasondoyle.ie](mailto:contact@jasondoyle.ie).

## References

1. OpenTelemetry, _What is OpenTelemetry?_, https://opentelemetry.io/docs/what-is-opentelemetry/.
2. OpenTelemetry, _Signals_, https://opentelemetry.io/docs/concepts/signals/.
3. CNCF TAG Observability, _Observability Whitepaper_, 2023, https://github.com/cncf/tag-observability/blob/main/whitepaper.md.
4. OpenTelemetry, _Collector_, https://opentelemetry.io/docs/collector/.
5. OpenTelemetry, _Sampling_, https://opentelemetry.io/docs/concepts/sampling/.
6. OpenTelemetry, _Metrics Data Model_, https://opentelemetry.io/docs/specs/otel/metrics/data-model/.
7. OpenTelemetry, _.NET metrics best practices_, https://opentelemetry.io/docs/languages/dotnet/metrics/best-practices/.
8. Prometheus, _Instrumentation_, https://prometheus.io/docs/practices/instrumentation/.
9. Prometheus, _Metric and label naming_, https://prometheus.io/docs/practices/naming/.
10. Grafana Labs, _4th annual Observability Survey_, https://grafana.com/observability-survey/.
11. Grafana Labs, _5 key takeaways from the Grafana Labs' 2024 Observability Survey_, 2024, https://grafana.com/blog/5-key-takeaways-from-the-grafana-labs-2024-observability-survey/.
12. New Relic, _2025 Observability Forecast press release_, 17 September 2025, https://newrelic.com/press-release/20250917.
13. Microsoft Learn, _Azure Monitor cost and usage_, archived source revision, 29 July 2025, https://github.com/MicrosoftDocs/azure-monitor-docs/blob/5b79f1a71a2523e964d70bbd7e9980e61d7d453a/articles/azure-monitor/fundamentals/cost-usage.md.
14. Google Cloud, _Google Cloud Observability pricing_, https://cloud.google.com/products/observability/pricing.
15. AWS, _Amazon CloudWatch pricing_, https://aws.amazon.com/cloudwatch/pricing/.
16. Rob Ewaschuk, _Monitoring Distributed Systems_, Google SRE, https://sre.google/sre-book/monitoring-distributed-systems/.
17. Vivek Rau, _Eliminating Toil_, Google SRE, https://sre.google/sre-book/eliminating-toil/.
18. Elizabeth Anna Mathilde Michels et al., _Alarm fatigue in healthcare_, BMC Nursing, 2025, https://doi.org/10.1186/s12912-025-03369-2.
19. Miriam Arnold, Mascha Goldschmitt, and Thomas Rigotti, _Dealing with information overload_, Frontiers in Psychology, 2023, https://doi.org/10.3389/fpsyg.2023.1122200.
20. ACM Digital Library, _Alert Fatigue in Security Operations Centres_, 2025, https://doi.org/10.1145/3723158.
21. OpenTelemetry, _Handling sensitive data_, https://opentelemetry.io/docs/security/handling-sensitive-data/.
22. European Union, _Article 5 GDPR_, https://gdpr-info.eu/art-5-gdpr/.
23. NIST, _SP 800-92: Guide to Computer Security Log Management_, 2006, https://csrc.nist.gov/pubs/sp/800/92/final.
24. CNCF, _How Uber is monitoring 4,000 microservices_, https://www.cncf.io/case-studies/uber/.
25. Uber Engineering, _M3: Uber's Open Source, Large-scale Metrics Platform for Prometheus_, 2018, https://www.uber.com/us/en/blog/m3/.
