Disclosure: These views are my own and do not represent my current or any former employers.
Executive summary
A model endpoint can be available, fast, and still be part of a failing product.
That is not a contradiction. It is a measurement error.
Traditional reliability engineering is very good at asking whether a service responds. Site reliability engineering uses service level indicators, service level objectives, and error budgets to align technical behaviour with user expectations. The usual indicators include availability, latency, error rate, throughput, and other properties of the serving system [1]. SRE also recognises that reliability is a product judgement, not the pursuit of perfect uptime at any cost [2].
Production AI adds a second layer. The endpoint may return a 200 response inside the latency budget while the product omits the critical fact, retrieves the wrong record, follows an injected instruction, sends an unauthorised message, hides a policy exception, or gives a plausible answer that cannot be replayed after the model changes. The service is up. The behaviour is wrong.
This paper argues for behavioural reliability engineering: the discipline required when software depends on probabilistic models, retrieval systems, prompts, policy instructions, tools, and external content. It does not replace infrastructure reliability. It extends it.
The public record supports the need for this distinction:
- NIST's AI Risk Management Framework and Generative AI Profile treat monitoring, measurement, documentation, information integrity, data provenance, security, and incident management as continuing lifecycle activities, not launch-time paperwork [3][4].
- Production machine learning literature has long warned that model behaviour can degrade because data, features, dependencies, and operating conditions change after deployment [5][6][7].
- Public LLM guidance and security work identify prompt injection, excessive agency, vector and embedding weaknesses, misinformation, and improper output handling as application risks that ordinary endpoint health checks do not capture [8][9][10][11].
- Public incidents and vendor documents show that AI products can mislead users, omit important details, expose protected context, or take destructive actions while the underlying platform remains reachable [12][13][14][15][16].
A credible production AI programme therefore needs more than uptime and p95 latency. It needs behavioural SLOs, evaluation sets tied to product obligations, omission and harmful-action rates, drift detection, retrieval quality controls, prompt-injection tests, deterministic policy boundaries, tool containment, provenance, replay packages, progressive delivery, rollback planning, and incident command that measures customer impact rather than endpoint health alone.
The strongest counterargument is practical. Teams already struggle to maintain basic reliability, and behavioural metrics can be expensive, subjective, and incomplete. Some product failures should be handled by user training, not by treating every generated error as an incident.
That counterargument is partly right. Behavioural reliability should not become a theatrical governance layer. It should be narrow, measurable, and attached to the decisions and actions that matter. The goal is to know whether the deployed product is safe enough for its intended use, whether it is getting worse, and what to do when it fails, rather than to prove that the model is generally good.
What matters is whether the product is still doing the job users and affected people are relying on it to do, not whether the endpoint is up.
1. The false pass
Most production dashboards can tell an operator whether the model gateway is responding. They can show request volume, token count, latency, server errors, rate limits, cost, and saturation. Those measures matter. A slow or unavailable endpoint can break a product immediately.
They are not sufficient.
Consider a support assistant that must identify refund eligibility. The endpoint returns in 900 milliseconds. The answer is fluent. The user accepts it. Only later does the organisation discover that a policy exception was omitted, a date was read from the wrong document, or a tool call changed the customer record before the precondition was checked.
From the model endpoint's point of view, the request succeeded. From the product's point of view, the work failed.
The same pattern appears in other workflows:
| Workflow | Endpoint result | Product failure |
|---|---|---|
| Inbox briefing | Fast summary returned | Critical message left out |
| Legal research | Answer returned | Non-existent authority cited |
| Retrieval assistant | Context retrieved | Wrong tenant, stale policy, or poisoned passage used |
| Coding agent | Task marked complete | Production data deleted or tests fabricated |
| Customer chatbot | Response produced | User receives incorrect entitlement or deadline |
| Security agent | Tool executed | Normal traffic blocked after false state inference |
Hallucination is only one failure mode among several: omission, misclassification, unsafe action, stale retrieval, hidden drift, bad tool scope, untraceable provenance, or a rollback that restores the code but not the customer harm.
SRE vocabulary is still useful here. An SLI should measure the level of service users actually experience [1]. If the service level of interest is correct handling of refund requests, surfacing of legal deadlines, or safe execution of tool calls, then availability and latency are proxy measures. They are necessary signals, not the reliability contract.
The false pass happens when the proxy becomes the claim.
2. Infrastructure reliability and behavioural reliability
Infrastructure reliability asks whether the system can serve requests under stated conditions. It includes availability, latency, durability, capacity, dependency health, deployment safety, and recovery.
Behavioural reliability asks whether the system's outputs and actions remain acceptable for the product's intended use under real operating conditions. It includes the model, prompt, retrieval layer, tools, policies, memory, context assembly, user interface, and human handoff.
The two layers interact but are not the same.
| Reliability question | Typical infrastructure measure | Behavioural equivalent |
|---|---|---|
| Is the service reachable? | Availability | Is the intended workflow completing correctly? |
| Is it fast enough? | Latency percentiles | Is verification effort still lower than doing the task manually? |
| Is it returning errors? | 5xx and timeout rate | Is it producing refusals, unsupported claims, or unsafe actions? |
| Did the release regress? | Canary error budget | Did golden, adversarial, and omission evals regress? |
| Can we debug it? | Logs, traces, metrics | Replay package with prompts, retrieved sources, tool calls, versions, and policy decisions |
| Can we recover? | Rollback, failover, restore | Undo, compensate, notify, quarantine memory, and retest behaviour |
NIST's AI RMF uses Govern, Map, Measure, and Manage as continuing functions for AI risk management [3]. The Generative AI Profile adds risks such as confabulation, data privacy, information integrity, information security, and value chain integration [4]. Those categories are not endpoint health metrics. They are application behaviour and lifecycle risks.
A product can pass the infrastructure layer and fail the behavioural layer. It can also pass a generic model benchmark and fail the product layer. A summarisation model may perform well on public datasets while omitting the one contractual exception this user's workflow depends on. A coding model may pass a benchmark while using a tool that has write access where only read access was required.
The boundary matters because it changes ownership. The model vendor may own endpoint availability and broad model safety work. The deploying organisation owns the product context: what sources are authoritative, what actions are allowed, what counts as harm, which users are affected, and how failure is detected.
3. Behavioural SLOs
A behavioural SLO is a measurable target for the product behaviour that matters to users or affected people.
It should be written as an operational claim, not as a mood. "The assistant should be helpful" is not an SLO. "In the billing-support evaluation set, at least 99.5 percent of high-consequence refund answers must select the correct policy path, cite the active policy source, and avoid tool execution unless all preconditions pass" is closer.
A behavioural SLO has five parts:
- Intended use.
- Population or workflow covered.
- Ground truth or evaluation method.
- Threshold and measurement window.
- Consequence when the threshold is missed.
Examples:
| Behavioural SLO | Candidate SLI |
|---|---|
| Critical obligations are surfaced | Critical-item recall over labelled cases |
| Summaries preserve material constraints | Mandatory-field omission rate |
| Answers are grounded | Share of material claims supported by approved sources |
| Retrieval uses current authorised sources | Retrieval precision, source freshness, and access-control pass rate |
| Tool actions are safe | Harmful-action rate and precondition failure rate |
| Handoffs occur when needed | Escalation recall for out-of-scope or high-risk cases |
| Drift is detected | Regression rate across stable eval cohorts |
| Incidents can be investigated | Replay package completeness rate |
The threshold should match consequence. A product recommending playlist names does not need the same omission budget as a system that prepares safety instructions, legal deadlines, customer refunds, or production changes. In high-impact action paths, the acceptable harmful-action rate may be zero unless a deterministic control prevents the action from executing.
This is where SRE error budgets need adaptation. An uptime error budget can tolerate a measured amount of unavailability because users can often retry. Some behavioural errors do not behave that way. A leaked document cannot be un-leaked. A message sent to a customer cannot be unsent. A medical contraindication omitted from a summary may matter even if the next request would have been correct.
That does not mean every behavioural SLO must be perfect, but the budget does need to be tied to reversibility, detectability, and harm. For low-impact drafts, an error budget can be broad. For irreversible tool actions, the SLO should be enforced by policy and permission boundaries, not by hope that the model usually chooses well.
4. Evaluation sets are product contracts
Public benchmarks are useful for comparing models. They rarely define whether a deployed product is reliable.
Production evaluation sets should be built from the product's obligations:
- common successful cases;
- edge cases that users still reasonably expect to work;
- high-consequence cases with strict escalation rules;
- adversarial inputs and indirect prompt injections;
- stale, conflicting, and missing-source cases;
- retrieval failures and access-control cases;
- tool-action cases with preconditions and rollback requirements;
- representative user language, including abbreviations and ambiguity;
- known incidents and near misses converted into regression tests.
The ML Test Score paper proposed production-readiness tests and monitoring needs for machine learning systems, including data, model, infrastructure, and monitoring checks [5]. Earlier work on hidden technical debt warned that ML systems create dependencies and feedback loops that are not obvious in ordinary code review [6]. A Microsoft case study of machine learning engineering described monitoring as part of the production workflow, not an optional final step [7].
Generative AI makes the same lesson more visible. OpenAI's Evals project states that without evaluations it can be difficult and time intensive to understand how different model versions affect a specific use case, and it supports custom and private evals for workflows that matter to the builder [17].
A good evaluation set is more than a scoring harness. It states what the organisation believes the product must not get wrong.
The scoring should separate failure types:
| Metric | What it reveals |
|---|---|
| Task success | Did the workflow reach the correct outcome? |
| Critical omission rate | Did the answer leave out material facts or constraints? |
| Unsupported-claim rate | Did generated claims lack approved evidence? |
| False refusal rate | Did the system fail to help when it should have? |
| Unsafe-compliance rate | Did it comply when it should have refused or escalated? |
| Retrieval miss rate | Did the retriever fail to fetch required context? |
| Retrieval contamination rate | Did irrelevant, stale, unauthorised, or poisoned material enter context? |
| Tool precondition pass rate | Were deterministic checks satisfied before action? |
| Harmful-action rate | Did the agent execute or propose actions outside policy? |
| Replay completeness | Can the case be reconstructed later? |
The hardest metric is omission. A false statement is visible in the output. A missing exception may leave no trace. That is why omission cases need labelled ground truth and explicit mandatory fields. If the product summarises customer obligations, the evaluation record should say which obligations must appear. If it prepares incident briefings, it should list the alerts, affected users, mitigation status, and open risks that cannot be left out.
5. Retrieval, omission, and drift
Many AI products are retrieval systems with a language model attached, not simply model calls.
Retrieval-augmented generation was introduced to combine parametric model knowledge with external sources for knowledge-intensive tasks [18]. In production, retrieval can fail in ways that look like model error:
- the correct source is absent from the index;
- the source is stale;
- the source exists but is ranked below irrelevant material;
- the source is in the middle of a long context and is underused;
- the source is retrieved from the wrong permission boundary;
- the source is poisoned, duplicated, or misleading;
- the source conflicts with a more authoritative record;
- the answer cites a source but the material claim is not supported by it.
RAGAS and related work evaluate dimensions such as context relevance, faithfulness, and answer relevance for RAG systems [19]. The "Lost in the Middle" study found that language models can perform worse when relevant information appears in the middle of long contexts, even when the information is present [20]. OWASP's 2025 Vector and Embedding Weaknesses category treats similarity search and embeddings as part of the application trust boundary, with risks including cross-tenant leakage, inversion, poisoning, and retrieval jamming [11].
The practical consequence is simple: do not measure only the answer. Measure the retrieval path.
A production RAG dashboard should show:
- source coverage by corpus and tenant;
- freshness of indexed material;
- access-control checks before and after retrieval;
- top-k recall for labelled queries;
- citation support for material claims;
- conflicts between retrieved sources;
- failed retrievals that led to confident answers;
- retrieval drift after embedding model, chunking, or ranking changes.
Drift is rarely confined to the model. User questions drift. Product policies drift. The indexed knowledge base drifts. The distribution of tool results drifts. A model version can remain unchanged while the product becomes less reliable because the environment around it changed.
That is why evaluation sets should have cohorts. Keep a stable regression set to detect change. Add recent production samples to detect new demand. Add incident-derived cases after every material failure. Measure each separately. If aggregate accuracy is flat but critical-new-sender recall or high-consequence refund handling falls, the product is deteriorating even though the average looks safe.
6. Prompt injection and agent tool use
Prompt injection is a structural problem created when instructions and data share the same context stream, not a strange edge case. OWASP's 2025 Prompt Injection entry describes direct and indirect inputs, retrieved content, tool output, multimodal content, memory, and agentic execution as delivery surfaces [9].
The risk grows when the model can act.
OWASP's Excessive Agency category identifies excessive functionality, excessive permissions, and excessive autonomy as root causes that allow damaging actions after unexpected, ambiguous, or manipulated model output [10]. The AI Agent Security Cheat Sheet recommends least privilege, per-tool permission scoping, explicit approval for sensitive operations, memory isolation, audit trails, rollback, and high-impact action controls [21].
The public record shows why this belongs in reliability as well as security. A prompt-injected assistant may leak data. It may also produce a false summary, skip a required approval, or call a tool with the wrong parameters. The user experiences a product failure.
Deterministic policy boundaries are therefore essential. The model may propose. It should not be the authority for actions that require policy judgement.
Practical boundaries include:
- read-only tools by default;
- separate identities for each tool and trust level;
- allowlisted operations instead of open-ended shell, browser, or database access;
- precondition checks outside the model;
- policy engines for spending, deletion, access, legal, safety, and external communication rules;
- idempotency keys and dry-run modes;
- rate, cost, and blast-radius limits;
- human approval bound to an exact action payload for high-impact steps;
- automatic quarantine of memory or retrieved content after suspected injection.
These controls are not a sign that the model is useless. They are how ordinary software engineering handles an untrusted component. OWASP's Misinformation entry makes the same operational point: incorrect or unsupported LLM output becomes a system-level failure when humans, automated workflows, or other agents act on it [22].
7. Public failures and what they prove
The public evidence is uneven. It is strong enough to reject endpoint uptime as a sufficient reliability claim. It is not strong enough to assign a universal failure rate to production AI.
Several cases are instructive.
Apple Intelligence produced inaccurate notification summaries that appeared to come from news organisations. The BBC reported false summaries involving criminal and sports news, and TechCrunch reported that Apple paused summaries for news and entertainment applications while it worked on changes [12][13]. The endpoint was not the central issue. The product created a misleading information surface.
In Moffatt v Air Canada, a Canadian tribunal held Air Canada responsible after a customer relied on incorrect bereavement-fare information supplied by the airline's chatbot [14]. The case is narrow, but the principle is useful: an automated interface is part of the service provider's product, not an independent actor outside responsibility.
In Mata v Avianca, lawyers submitted non-existent judicial opinions and fake quotations generated through ChatGPT. The court imposed sanctions and emphasised counsel's gatekeeping duty [15]. That was not a production endpoint outage. It was overreliance on a generated artifact in a consequential workflow.
Microsoft documented CVE-2025-32711 as an M365 Copilot information disclosure vulnerability [16]. The EchoLeak research described a zero-click prompt-injection exploit path against a production LLM system, with important caveats about responsible disclosure and mitigation [23]. These examples show that prompt injection and context handling can be product security and reliability issues.
Replit publicly discussed an incident in which a user's app database was deleted during agent use. Replit said changes were backed up and recoverable, but also said the user experience was bad because the agent was unaware of rollback and because development actions could affect production before its newer development and production database separation shipped [24]. Coding agents should not avoid data entirely; rather, recovery features, environment separation, and agent knowledge need to be part of the product's reliability.
Microsoft's own documentation for Outlook's Prioritize feature states that some mail is not evaluated, including mail delivered outside the Inbox, meeting mail, some encrypted messages, and other categories [25]. Its email summary FAQ says summaries may overlook important details or misinterpret context and advises users to review original content for critical information [26]. The UK Government Digital Service trial of M365 Copilot reported positive time savings and user sentiment, while also noting limitations for complex, nuanced, or data-heavy work and concerns about security and sensitive data in some cases [27].
The evidence supports a balanced conclusion. AI products can provide real benefit. They can also fail in ways that are not captured by endpoint health.
8. Provenance and replay packages
A conventional incident review asks what changed. In AI products, that question can be hard to answer unless provenance is designed in advance.
The relevant change might be:
- model vendor or model version;
- decoding parameters;
- system prompt;
- prompt template;
- user instructions;
- safety policy;
- retrieval corpus;
- embedding model;
- chunking or ranking configuration;
- tool schema;
- permission grant;
- feature flag;
- memory content;
- user interface wording;
- evaluation grader;
- post-processing rule.
Model Cards and Datasheets for Datasets proposed structured reporting for models and datasets before the current wave of LLM products [28][29]. The same idea should be operationalised at runtime.
A replay package is the minimum record needed to reconstruct a consequential case after the product changes. It should contain:
- request identifier, time, tenant, product surface, and user role;
- model provider, model identifier, model version if available, and parameters;
- prompt template version and system or developer instruction hash;
- user prompt and relevant conversation state, subject to privacy controls;
- retrieved source identifiers, versions, timestamps, hashes, and access decisions;
- tool schemas available to the model;
- tool calls requested, approved, denied, and executed;
- tool outputs and receipts;
- policy engine decisions;
- feature flag and configuration state;
- final output shown to the user;
- human approval record, if any;
- evaluation result, if the case was replayed later;
- redaction record for sensitive data.
This does not mean every prompt should be stored forever in raw form. Privacy, retention, and security rules still apply. Some fields may need hashing, encryption, sampling, or tenant-specific retention.
The point is that an incident cannot depend on memory and screenshots. If a customer asks why an agent sent a message, an operator should be able to identify the exact model, prompt, retrieved sources, tool payload, policy decision, and approval boundary that produced it.
9. Progressive delivery and rollback limits
Progressive delivery is valuable for AI systems, but only if the canary measures the behaviour that matters.
Feature flags and canaries are well-established release techniques. Martin Fowler notes that feature-flagged systems should expose their current toggle configuration and that flags introduce validation complexity because multiple code paths may be live [30]. The SRE Workbook defines canarying as partial, time-limited deployment of a change and evaluation against a control before broader rollout [31].
For AI products, the change may be a model, prompt, retrieval index, tool permission, grounding policy, memory strategy, or UI copy. A useful canary should compare both infrastructure and behaviour:
- endpoint availability and latency;
- cost and token use;
- task success;
- critical omission rate;
- escalation rate;
- unsupported claims;
- retrieval freshness;
- tool denial and approval rates;
- harmful-action attempts blocked by deterministic controls;
- customer correction and complaint signals;
- human verification time.
Rollback is necessary but not sufficient.
Rolling back a prompt does not retract a message already sent. Rolling back a model does not remove a hallucinated answer from a customer's decision process. Restoring a database does not undo the time a production system was broken. Removing poisoned content from an index may not remove it from memory, logs, embeddings, downstream summaries, or user exports.
Recovery planning should therefore distinguish four cases:
- Reversible behaviour: regenerate, correct, or reroute.
- Compensable behaviour: notify, apologise, refund, repair, or provide a corrected record.
- Containable behaviour: stop further spread by disabling tools, quarantining memory, or removing a source.
- Irreversible behaviour: treat as an incident with customer impact even if the endpoint stayed healthy.
A rollback plan that only says "revert to the previous model" is one action in a wider incident response, not a complete recovery plan.
10. Incident command and customer impact
An AI behaviour incident should be declared when the product materially misleads, omits, discloses, or acts outside its intended bounds, even if ordinary uptime dashboards are green.
Google's incident management guidance emphasises clear roles, including an incident commander, operations lead, planning, and communications, to avoid uncoordinated changes during incidents [32]. The same structure applies to AI behaviour incidents, but the roles need additional expertise.
A practical incident team may include:
- incident commander;
- operations lead for feature flags, kill switches, and deployment state;
- behaviour lead for evals, prompts, model versions, and failure taxonomy;
- retrieval lead for indexes, sources, permissions, and freshness;
- tool-safety lead for action logs and containment;
- security or privacy lead where disclosure or injection is involved;
- customer-impact lead;
- communications lead;
- provenance custodian to preserve replay packages;
- product owner to decide whether intended use must change.
Severity should include customer impact as well as system impairment:
| Severity signal | Example |
|---|---|
| Harmful action | Agent changed, deleted, sent, paid, blocked, or approved outside policy |
| Critical omission | System failed to surface a legal, safety, financial, or security obligation |
| Misleading authority | Output appeared to come from a trusted source but contradicted it |
| Data disclosure | Model, retrieval, tool, log, or output exposed unauthorised information |
| Containment failure | Poisoned memory or retrieved content may affect future sessions |
| Investigation failure | No replay package exists for affected cases |
Customer impact work must not wait until root cause is known. If the affected set can be bounded by logs or replay packages, bound it. If it cannot, say so internally and treat the uncertainty as part of the incident. A low number of confirmed complaints is not proof of low impact when the failure mode is omission.
Postmortems should avoid blaming the model as if it were a person. The useful questions are about system design:
- Why was the model able to act without the missing precondition?
- Why did the evaluation set not include this case?
- Why was the source stale or unauthorised?
- Why was the omission not observable?
- Why was there no replay package?
- Why did rollback not address customer harm?
SRE postmortem practice values learning from failure and acting on findings [33]. AI incidents should do the same.
11. A practical operating model
A workable operating model should be small enough to use and strict enough to matter.
Classify the workflow
Place each AI workflow into a risk tier:
| Tier | Example | Default control |
|---|---|---|
| Advisory | Drafting, brainstorming, low-impact summaries | User review and lightweight evals |
| Decision support | Customer support, policy interpretation, incident briefing | Behavioural SLOs, grounding, escalation, replay |
| Bounded action | Updating tickets, preparing records, internal workflow changes | Deterministic preconditions, action preview, limited tools |
| High-impact action | External messages, financial changes, deletion, access changes | Human approval bound to exact payload, strong policy engine |
| Prohibited autonomy | Legal commitment, safety-critical decision, irreversible destructive action | Model may not execute; use deterministic process |
Define the product contract
For each workflow, write:
- intended use;
- prohibited reliance;
- authoritative sources;
- material omissions;
- allowed actions;
- escalation rules;
- SLOs and thresholds;
- named owner;
- incident triggers.
Build evals before launch
At minimum, include:
- golden cases;
- high-consequence cases;
- refusal and escalation cases;
- prompt-injection cases;
- retrieval-miss and stale-source cases;
- tool precondition cases;
- regression cases from incidents;
- cases with expected omissions explicitly labelled.
Monitor behaviour in production
Use sampling, user feedback, shadow evaluation, automated graders where appropriate, and human review for high-consequence cohorts. Do not collapse the result into one accuracy number. Track the failure modes separately.
Enforce policy outside the model
The model should not be the only thing preventing data deletion, customer communication, entitlement approval, access grant, or payment. Put deterministic checks where ordinary software can enforce them.
Require replay for consequential work
If the organisation would need to explain the result to a customer, regulator, court, auditor, or incident review, the system should preserve enough provenance to replay or at least reconstruct the case.
Run behavioural release reviews
Before changing a model, prompt, retriever, tool, or policy, review:
- which behavioural SLOs are affected;
- which eval sets passed or failed;
- which cohorts are excluded;
- canary plan;
- rollback and compensation plan;
- monitoring window;
- named decision maker.
This is release engineering for systems whose failures are behavioural, not generic AI ethics prose.
12. The strongest counterargument
The strongest counterargument is that the paper asks teams to build a parallel reliability discipline before the evidence and tooling are mature.
There are several concerns.
First, behavioural metrics can be expensive. Labelling omissions, evaluating groundedness, and reviewing tool actions take time. Smaller teams may not have the capacity.
Second, many AI products are assistive. If a user is told to verify the output, perhaps the responsibility should remain with the user. Over-instrumenting the system may create false confidence and slow useful adoption.
Third, some failures are not unique to AI. Search systems return stale results. Humans omit details. Rule-based systems misroute work. A special AI reliability programme could become compliance theatre.
Fourth, the public evidence is incomplete. There is no comprehensive incident register for production AI failures. Vendors do not routinely publish critical omission rates, harmful-action rates, prompt-injection success rates, or false-negative rates for product-specific workflows. Peer-reviewed evidence on agentic systems in production remains limited.
These objections should change the design, not remove the need.
The response is to scope behavioural reliability to material workflows. Do not create a heavyweight process for every autocomplete feature. Do create measurable controls where the product allocates attention, interprets policy, handles sensitive data, or acts through tools.
User verification is also not a complete answer. Verification changes the product's value. If the user must re-open every source and inspect every tool action, the measured benefit becomes the net benefit after review, correction, and recovery, not the time taken to generate the answer.
Finally, the fact that humans and ordinary software also fail does not excuse ignoring new failure modes. Familiar reliability ideas still apply: define the service level, measure it, set limits, release gradually, contain blast radius, investigate incidents, and improve the system.
13. What this paper does not claim
This paper does not claim that endpoint availability and latency are unimportant. They remain necessary production measures.
It does not claim that every AI product is high risk. Many uses are low consequence and should have proportionate controls.
It does not claim that all model errors are incidents. An incident depends on intended use, impact, scope, and recoverability.
It does not claim that generic benchmarks are useless. They are useful evidence about model capability, but they are not a substitute for product-specific evaluation.
It does not claim that deterministic rules are perfect. Rules can be wrong, stale, or incomplete. Their value is that policy-critical boundaries can be inspected, tested, and enforced outside a probabilistic model.
It does not claim that human review automatically solves the problem. Human review must be attached to clear evidence, exact action payloads, and realistic time for judgement.
It does not claim that the public incidents cited here prove a single industry-wide failure rate. The evidence base is incomplete, and some cases are legal disputes, vendor disclosures, controlled demonstrations, or public reports rather than peer-reviewed incident datasets.
It does claim this:
If an AI product can omit, mislead, disclose, or act while its endpoint remains available, then reliability engineering must measure and control those behaviours. Uptime is not the product contract.
Conclusion
Production AI has inherited the language of reliable services without yet adopting the full implication of that language.
An SLO is supposed to describe what matters to the user. For many AI products, what matters is whether the right sources were considered, whether material facts were preserved, whether the answer was grounded, whether the action was authorised, whether the behaviour changed after a model or prompt update, and whether the organisation can reconstruct what happened when it fails. Whether a request merely returned is not the same question.
The endpoint can be up while the product is failing.
That failure may be quiet. It may be an omitted obligation, a stale retrieval result, an over-permissive tool, a poisoned memory, a false state inference, or a customer who trusted an automated answer because the interface presented it as part of the service.
The answer is to engineer the boundaries, rather than to freeze AI systems or require a human in every loop:
- behavioural SLOs for consequential workflows;
- evaluation sets that encode product obligations;
- omission and harmful-action metrics;
- retrieval and drift monitoring;
- prompt-injection and agent tests;
- deterministic policy enforcement;
- provenance and replay;
- progressive delivery;
- recovery plans that address customer harm;
- incident command that can operate when the dashboard is green.
The reliability question for production AI is not whether the model endpoint is alive.
It is whether the product remains worthy of the reliance it invites.
About the author
Jason Doyle writes about reliable software, observability, applied AI, and practical controls for systems that influence human decisions. He publishes at jasondoyle.ie and can be contacted at contact@jasondoyle.ie.
References
- Google Site Reliability Engineering, Service Level Objectives, https://sre.google/sre-book/service-level-objectives/.
- Google Site Reliability Engineering, Embracing Risk, https://sre.google/sre-book/embracing-risk/.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023, https://doi.org/10.6028/NIST.AI.100-1.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024, https://doi.org/10.6028/NIST.AI.600-1.
- Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley, The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction, IEEE Big Data, 2017, https://research.google/pubs/the-ml-test-score-a-rubric-for-ml-production-readiness-and-technical-debt-reduction/.
- D. Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS, 2015, https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-systems.
- Saleema Amershi et al., Software Engineering for Machine Learning: A Case Study, ICSE, 2019, https://www.microsoft.com/en-us/research/publication/software-engineering-for-machine-learning-a-case-study/.
- OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2025, released 18 November 2024, https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf.
- OWASP GenAI Security Project, LLM01:2025 Prompt Injection, in OWASP Top 10 for LLM Applications 2025, https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf.
- OWASP GenAI Security Project, LLM06:2025 Excessive Agency, in OWASP Top 10 for LLM Applications 2025, https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf.
- OWASP GenAI Security Project, LLM08:2025 Vector and Embedding Weaknesses, in OWASP Top 10 for LLM Applications 2025, https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf.
- BBC News, Apple AI notification errors persist despite complaints, 6 January 2025, https://www.bbc.co.uk/news/articles/cge93de21n0o.
- TechCrunch, Apple pauses AI notification summaries for news after generating false alerts, 16 January 2025, https://techcrunch.com/2025/01/16/apple-pauses-ai-notification-summaries-for-news-after-generating-false-alerts/.
- Moffatt v Air Canada, 2024 BCCRT 149, https://canlii.ca/t/7np8t.
- Mata v Avianca, Inc., No. 1:22-cv-01461, Opinion and Order on Sanctions, 22 June 2023, https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2022cv01461/575368/54/.
- Microsoft Security Response Center, CVE-2025-32711: M365 Copilot Information Disclosure Vulnerability, 11 June 2025, https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711.
- OpenAI, Evals, https://github.com/openai/evals.
- Patrick Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, NeurIPS, 2020, https://arxiv.org/abs/2005.11401.
- Shahul Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation, EACL Demo, 2024, https://aclanthology.org/2024.eacl-demo.16/.
- Nelson F. Liu et al., Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, 2024, https://aclanthology.org/2024.tacl-1.9/.
- OWASP Cheat Sheet Series, AI Agent Security Cheat Sheet, https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html.
- OWASP GenAI Security Project, LLM09:2025 Misinformation, in OWASP Top 10 for LLM Applications 2025, https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf.
- Pavan Reddy and Aditya Sanjay Gujral, EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System, 2025, https://arxiv.org/abs/2509.10540.
- Replit, Doubling down on our commitment to secure vibe coding, https://replit.com/blog/doubling-down-on-our-commitment-to-secure-vibe-coding.
- Microsoft Support, Prioritize my inbox, https://support.microsoft.com/en-us/outlook/copilot-outlook/prioritize-my-inbox.
- Microsoft Learn, FAQ for email summary feature in Outlook, updated 13 May 2026, https://learn.microsoft.com/en-us/microsoft-sales-copilot/faqs-email-summary.
- UK Government Digital Service, Microsoft 365 Copilot Experiment: Cross-Government Findings Report, 2025, https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report/microsoft-365-copilot-experiment-cross-government-findings-report-html.
- Margaret Mitchell et al., Model Cards for Model Reporting, FAT*, 2019, https://arxiv.org/abs/1810.03993.
- Timnit Gebru et al., Datasheets for Datasets, Communications of the ACM, 2021, https://arxiv.org/abs/1803.09010.
- Pete Hodgson, Feature Toggles, Martin Fowler, 2017, https://martinfowler.com/articles/feature-toggles.html.
- Google Site Reliability Engineering Workbook, Canarying Releases, https://sre.google/workbook/canarying-releases/.
- Google Site Reliability Engineering, Managing Incidents, https://sre.google/sre-book/managing-incidents/.
- Google Site Reliability Engineering Workbook, Postmortem Culture: Learning from Failure, https://sre.google/workbook/postmortem-culture/.