Disclosure: These views are my own and do not represent my current or any former employers.
Executive summary
Incident command is often described as a way to co-ordinate responders during an outage. That is true, but incomplete. The deeper function is architectural.
Incident command is the interface between a failing technical system and the organisation trying to recover it.
A service outage, data incident, security compromise, or failed change does not present itself to the organisation as a neat problem statement. It arrives as symptoms, alerts, partial telemetry, customer reports, internal speculation, unclear authority, legal constraints, anxious executives, vendor dependencies, and people under pressure. The incident command process turns that disorder into a recoverable operating model.
The speed of recovery therefore depends less on heroic debugging than on how four things move through the interface:
- Evidence: what is known, how it is trusted, where uncertainty remains, and who can see it.
- Authority: who may decide, who may act, and which decisions require escalation.
- Ownership: which person or team owns each thread of work until it is closed.
- Communication: how responders, customers, executives, vendors, regulators, and support teams receive accurate status without interrupting the work.
Public guidance supports this view. NIST SP 800-61 Revision 3 maps incident response into the Cybersecurity Framework 2.0 functions of Govern, Identify, Protect, Detect, Respond, and Recover, making incident response part of risk management rather than a purely technical workflow.[1][2] CISA's federal incident and vulnerability response playbooks emphasise standard procedures to identify, co-ordinate, remediate, recover, and track mitigations across organisations.[3] Google's SRE material describes incident management as a way to co-ordinate, communicate, and control response under stress, with explicit roles such as Incident Commander, Operations Lead, Communications Lead, and Planning Lead.[6]
Public incidents show the same pattern. AWS S3 in 2017, Facebook in 2021, Atlassian in 2022, and CrowdStrike in 2024 all involved more than a technical defect. Each account also exposed dependencies, recovery constraints, communication needs, or cross-organisation coordination.[10][11][12][13][14]
These are not stories about one engineer being insufficiently clever. They are stories about interfaces. The organisation had to discover the state of the system, decide what mattered, allocate authority, prevent uncoordinated change, inform people who were not in the room, recover safely, and learn without hiding the uncomfortable parts.
This paper proposes a concrete model that can be measured. During and after an incident, the organisation should be able to identify the evidence that changed the response, the authority behind material actions, the owner of each workstream, the freshness of stakeholder updates, the mitigations taken before root cause was known, and the handoffs that preserved or lost context.
The claim is not that process fixes incidents. Bad process can slow recovery, suppress local expertise, and create a false sense of control. The claim is narrower and more practical: when a technical system is failing, an organisation needs an interface that allows the right people to see, decide, own, and communicate the right things at the right time. Incident command is that interface.
1. The interface between system failure and organisational action
A technical system can fail faster than an organisation can understand it.
A cache starts returning stale data. A region loses capacity. A security tool delivers a harmful update. A database replica falls behind. Customers report behaviour that telemetry does not yet explain. A dependent provider degrades. A feature flag, migration, network route, certificate, queue, or access policy behaves in a way that was individually plausible and collectively surprising.
The first organisational problem is not root cause. It is translation.
The system expresses failure through signals:
- alerts;
- logs;
- traces;
- metrics;
- exceptions;
- synthetic checks;
- customer tickets;
- support calls;
- social media reports;
- partner notifications;
- dashboards that may themselves be degraded.
The organisation must convert those signals into action:
- declare or decline an incident;
- set severity;
- assign roles;
- protect customers;
- select mitigations;
- stop unsafe change;
- decide when to escalate;
- contact vendors or authorities;
- prepare customer and executive communication;
- recover data or service;
- record what happened;
- learn afterwards.
Incident command is the structure that makes this conversion possible.
This is why an incident cannot be treated only as a debugging problem. Debugging asks, "What is wrong?" Incident command also asks:
- What do we know now?
- What are we assuming?
- What is the safest next action?
- Who can make that action?
- What is the cost of waiting?
- What is the cost of being wrong?
- Who needs to know before the next update?
- What must be preserved for later investigation?
The difference matters because recovery work often begins before diagnosis is complete. NIST's current incident-response publication situates response and recovery inside broader governance, asset understanding, protection, detection, and improvement.[1] That framing is important. An incident is best understood as a temporary operating mode for the organisation, rather than a ticket owned by one technical team.
The four channels
A concrete incident-command interface has four channels.
| Channel | Question | Typical failure | Useful measure |
|---|---|---|---|
| Evidence | What is true enough to act on? | Teams debate impressions while telemetry, customer impact, and assumptions are mixed together | Time from signal to verified incident fact; percentage of key claims with source links |
| Authority | Who can decide or act? | Engineers wait for permission or take conflicting action without permission | Time to decision for material mitigations; number of actions outside command |
| Ownership | Who owns each thread? | Important work is everybody's concern and nobody's task | Percentage of open workstreams with named owner and next check-in |
| Communication | Who needs to know what, when? | Responders are interrupted for status while customers and leaders receive stale or speculative updates | Update freshness by audience; number of stakeholder interruptions to operators |
This model is deliberately simple. A five-person startup and a multinational platform company both need evidence, authority, ownership, and communication. The difference is scale.
The interface also has direction. Evidence must move up from the system to command. Decisions must move down from command to operators. Customer signals must move sideways into diagnosis. Constraints must move from legal, security, privacy, support, and business teams into the response. Status must move out without forcing every responder to explain the same facts repeatedly.
When those flows are healthy, the incident can still be hard, but it becomes governable. When they are blocked, even talented people can make the incident worse.
2. Command roles are not status titles
Incident roles are not badges. They are decision surfaces.
Google's incident-management guidance describes an incident response system based on the Incident Command System and organised around co-ordination, communication, and control. Its core roles include Incident Commander, Communications Lead, and Operations Lead, with the ability to delegate work as the incident grows.[6] The SRE book's incident chapter adds that the Incident Commander holds high-level state, assigns responsibilities, and keeps the living incident document; the Operations Lead applies operational tools; the Communications role sends periodic stakeholder updates; and Planning handles longer-term issues, handoffs, and tracking how the system has diverged from normal.[7]
The important feature is separation of concerns.
An on-call engineer deep in a log stream should not also decide company-wide status language, handle executive questions, approve risky changes, co-ordinate vendor escalation, preserve evidence, and plan the next handoff. That overloads the person closest to the failure and makes the organisation dependent on whoever is staring hardest at the terminal.
A useful command model separates at least five functions:
| Function | Purpose | Typical owner |
|---|---|---|
| Command | Maintain overall state, set priorities, assign work, make or route decisions | Incident Commander |
| Operations | Mitigate customer impact and make controlled technical changes | Operations Lead and delegated operators |
| Investigation | Test hypotheses, collect evidence, and narrow uncertainty | Diagnosis lead or investigation owners |
| Communications | Produce internal, executive, customer, support, and partner updates | Communications Lead |
| Planning and recovery | Track next steps, handoffs, dependency restoration, and return to normal | Planning Lead or recovery owner |
Small incidents can combine roles. One person may hold command and communications for a minor issue. The interface should still make the roles visible. If nobody is explicitly handling communication, the on-call engineer will handle it by interruption. If nobody owns planning, recovery tasks become memory. If nobody owns command, authority will be improvised by whoever speaks first or loudest.
Decision rights
Incident command works only if decision rights are explicit before the incident.
The following questions should not be invented at three in the morning:
- Who may declare severity one?
- Who may take a region, tenant, feature, or data pipeline out of service?
- Who may roll back a release?
- Who may disable a security control temporarily?
- Who may trigger customer notification?
- Who may invoke disaster recovery?
- Who may accept degraded operation rather than continued investigation?
- Who may speak to regulators, media, or major customers?
- Who may close the incident?
NIST's 2025 revision makes governance part of the incident-response profile, including roles, responsibilities, policies, and risk context.[1] That is precompiled authority rather than bureaucracy for its own sake. During an incident, ambiguity about authority is latency.
Authority should be narrower than accountability. An Incident Commander does not need to know every implementation detail. They do need authority to choose priorities, stop unsafe work, request evidence, escalate for business decisions, and assign ownership. A technical owner may have authority to execute a rollback but not to decide whether the public status page should describe customer impact as resolved. Legal or privacy teams may own notification obligations but should not become a bottleneck for an urgent containment action already allowed by policy.
The interface should encode these boundaries in runbooks, access controls, deployment tooling, and incident templates. A sentence in a playbook is not enough if the production system allows ten people to make conflicting changes.
3. Mitigation and diagnosis are different jobs
A common incident failure is to confuse knowing the cause with reducing harm.
Diagnosis seeks explanation. Mitigation seeks lower impact. Recovery often requires both, but not in sequence.
If a new release appears correlated with errors, rollback may be the right mitigation before the exact faulty line is known. If a region is unhealthy, traffic may need to be drained before the cause is found. If a credential is exposed, revocation may precede full investigation. If a data migration is corrupting records, stopping writes may matter more than explaining the corruption mechanism.
Google's SRE incident examples repeatedly make this distinction practical. Its emergency-response chapter describes incidents where monitoring, rollback paths, out-of-band communication, and tested procedures mattered because responders needed to act under uncertainty.[8] The incident-management guide explicitly states that fixing the problem is only part of response; users, stakeholders, and leaders also need accurate information about impact, workarounds, mitigation, and resolution.[6]
The Incident Commander should therefore keep two boards, even if they live in one document.
| Board | Question | Examples |
|---|---|---|
| Mitigation board | How do we reduce user, data, safety, security, or business impact now? | Roll back, drain traffic, disable feature, revoke key, fail over, pause batch job, increase capacity |
| Diagnosis board | What evidence explains the failure and prevents unsafe recovery? | Compare deploys, inspect traces, query logs, reproduce condition, check dependency status, examine audit trail |
The boards interact. A mitigation can create evidence. A diagnostic result can change the mitigation. But separating them prevents a familiar trap: the team spends thirty minutes arguing about the root cause while a simple reversible action could have reduced customer harm.
It also prevents the opposite trap: the team changes the system repeatedly without understanding whether each change helped. The AWS S3 2017 report is useful here. An authorised team member used an established playbook, but an incorrect input removed a larger set of servers than intended. The eventual changes included tool safeguards, slower capacity removal, minimum-capacity checks, audits of other operational tools, and further partitioning to reduce blast radius.[10] Operators still needed to act. What was missing was guardrails around that action and evidence about its effect.
Evidence flow
Evidence is not whatever appears in the loudest channel.
During an incident, evidence should be labelled by quality:
| Evidence class | Example | Command handling |
|---|---|---|
| Observed fact | Error budget burn, request failure rate, affected customer count, specific alert, command audit record | Record source and timestamp |
| Correlation | Error increase began after deployment, support tickets started after provider alert | Use for hypothesis and mitigation, not final cause |
| Hypothesis | Cache invalidation loop, bad config, expired certificate, upstream limit | Assign an owner and test |
| Decision | Roll back version X, pause job Y, notify customers Z | Record approver, time, expected effect |
| Constraint | Do not delete evidence, preserve logs, regulatory notice deadline, customer embargo | Put at top of incident document |
| Unknown | Scope unclear, data impact not yet measured, vendor ETA absent | Communicate as unknown rather than filling the gap with confidence |
This structure reduces argument. It lets a customer-facing update say, "We have confirmed increased error rates in service A since 10:12 UTC. We have not yet confirmed data loss." It lets operations say, "Rollback reduced five hundred errors per minute to forty, but latency remains elevated." It lets executives see uncertainty without receiving a detective story.
Evidence flow also needs provenance. Dashboards should link to queries. Customer-impact estimates should say which source produced them. Change correlations should link to deployment IDs. Vendor reports should identify the provider and time received. Without provenance, the incident document becomes a rumour with headings.
4. Span of control and the shape of escalation
Incidents grow by attracting helpers.
That can be good. A major incident may require storage, networking, application, security, support, legal, product, executive, vendor, and customer-success participation. It may require twenty people, not three.
The problem is that attention does not scale linearly. FEMA's Incident Command System guidance defines manageable span of control as three to seven people or resources for one supervisor, with five considered optimal.[5] A commander directly co-ordinating fifteen workstreams is not commanding. They are switching context until something important falls out.
Technical incidents need the same compression.
Instead of inviting every expert into one call and asking each to speak whenever they discover something, the command structure should create workstreams:
- customer impact;
- recent change analysis;
- infrastructure capacity;
- dependency and vendor status;
- security and evidence preservation;
- data integrity;
- mitigation execution;
- customer communication;
- executive and regulatory coordination;
- recovery validation.
Each workstream has one owner. Owners can have their own helpers. The Incident Commander receives concise state from owners, not a live transcript of everyone thinking.
This protects cognition rather than enforcing hierarchy for its own sake. It keeps the command interface narrow enough to operate.
Handoffs
Long incidents require handoffs. Bad handoffs are incident multipliers.
The SRE book recommends explicit handoff of the Incident Commander role, including direct acknowledgement from the incoming commander and communication to the wider incident group.[7] That practice matters because authority must not become ambiguous during shift change.
A useful handoff contains:
- current severity and customer impact;
- confirmed facts with timestamps;
- active mitigations and their observed effect;
- open hypotheses and owners;
- decisions already made and why;
- actions explicitly rejected;
- current constraints;
- stakeholder commitments and next update times;
- unresolved risks;
- where to find the incident log and evidence.
The outgoing commander should not hand off a mood. They should hand off the interface state.
Fatigue changes this from good practice to safety requirement. As an incident extends, the people with the best context become the people most likely to be tired, hungry, tunnelled, and reluctant to leave. Planning needs to assign relief before judgement degrades. The Planning role in Google's model explicitly includes handoffs and logistical needs.[7] Food, sleep, timezone coverage, and relief are not soft concerns. They preserve decision quality.
5. Multi-service and multi-organisation incidents
Modern incidents frequently cross boundaries that the organisation chart treats as separate. Storage failures affect compute and dashboards. Security updates affect customer endpoints. Vulnerabilities cross open-source and commercial stacks. SaaS recovery depends on cloud primitives, orchestration, support, and customer validation.
The command interface must handle dependencies as well as tasks.
Public examples
In the 2017 Amazon S3 incident, the initial capacity-removal error affected the index and placement subsystems. S3 APIs became unavailable, and other AWS services in US-EAST-1 that relied on S3 were affected. The AWS Service Health Dashboard administration console also depended on S3, so AWS used Twitter and banner text until dashboard updates were available.[10] The technical failure and the communication interface shared a dependency.
In Facebook's 2021 outage, a command intended to assess backbone capacity unintentionally took down backbone connections. DNS servers withdrew BGP advertisements because they could not communicate with data centres, making Facebook services unreachable. Normal internal tools and access paths were also broken, so engineers needed onsite access under high physical-security constraints.[11] Recovery depended on understanding both the service and the access environment.
Atlassian's April 2022 outage affected approximately 400 cloud customers after a team-to-team communication gap and script execution error led to improper deletion of sites. Restoration was slow because affected customer data had to be extracted and restored into an active multi-tenant environment, with internal validation and customer verification.[12] The interface had to manage technical restoration and customer-by-customer trust together.
The CrowdStrike Falcon content update in July 2024 illustrates multi-organisation recovery. CrowdStrike stated that a problematic Rapid Response Content configuration update caused Windows crashes for in-scope hosts and that the update was reverted. Microsoft wrote that the event was not a Microsoft incident but affected the Windows ecosystem, estimated 8.5 million Windows devices were affected, and described collaboration among Microsoft, CrowdStrike, AWS, GCP, customers, and external developers.[13][14] Command did not belong to a single organisation. Each affected organisation still needed local command to triage endpoints, prioritise critical services, communicate internally, and coordinate with external guidance.
The Log4j vulnerability response showed a different form of boundary crossing. The Cyber Safety Review Board described the vulnerability as serious and endemic, with response challenges involving asset discovery, vendor coordination, open-source dependencies, and long-tail remediation.[16] In this class of incident, the question "Are we affected?" can be harder than applying the patch.
These examples differ in cause and scale. They support a common claim: incident command must represent the actual dependency graph, not the ownership diagram.
The dependency register
For major incidents, the incident document should contain a live dependency register.
| Dependency | Owner | Status | Evidence | Next action | Escalation path |
|---|---|---|---|---|---|
| Cloud provider region | Platform owner | Degraded | Provider status and synthetic checks | Test failover | TAM or support case |
| Identity provider | Security owner | Unknown | Login failures, audit logs incomplete | Confirm scope | Vendor incident contact |
| Customer support queue | Support lead | Backlogged | Ticket count and top categories | Publish macro | Support director |
| Data restore | Database owner | In progress | Restore job IDs | Validate tenant A | Engineering director |
This table seems mundane until the incident grows. Then it becomes the difference between co-ordination and hope.
6. Customer and executive communication
Communication is part of recovery work itself, carried out alongside engineering rather than added afterwards as a courtesy.
Customers need to know whether they are affected, what they should do, what the provider is doing, when the next update will arrive, and what remains unknown. Support teams need consistent language. Executives need enough detail to make business decisions without dragging responders into repeated briefings. Legal and privacy teams need facts early enough to assess notification obligations. Regulators may require timely reporting. Public companies in the United States, for example, have cybersecurity disclosure obligations for material incidents under SEC rules, with disclosure generally due within four business days after a materiality determination.[22] CISA also asks organisations reporting incidents to provide as much detail as possible to support timely handling and analysis.[4]
The Communications Lead should not be a stenographer. They translate the incident state for different audiences while protecting responders from interruption.
| Audience | Needs | Common mistake |
|---|---|---|
| Responders | Current state, decisions, owners, next check-in | Mixing executive narrative into technical channel |
| Support | Customer-visible symptoms, affected products, approved workarounds, escalation instructions | Asking support to infer from engineering chat |
| Customers | Impact, mitigation, workarounds, next update, apology when warranted | Overstating certainty or hiding known impact |
| Executives | Impact, risk, decision requests, customer exposure, regulatory triggers | Demanding root cause before mitigation |
| Vendors | Technical evidence, timestamps, account IDs, severity | Escalating without a clear ask |
| Regulators or authorities | Required facts, legal classification, timing, contact | Letting public updates outrun confirmed obligations |
Google's incident-management guide states that users, stakeholders, and leaders need updates about what is affected, severity, workarounds, mitigation, and resolution, and that consistent communication builds trust and transparency.[6] The AWS S3 report's discussion of the Service Health Dashboard dependency shows why communication channels must be resilient to the incident itself.[10]
A practical communication rule is simple: if a status has a consumer, it needs an owner and a clock.
- Internal responder update every fifteen to thirty minutes for active major incidents.
- Executive update at a cadence matched to decision need.
- Customer update at a published cadence, even when the update is "no material change".
- Support macro updated whenever customer-facing facts change.
- Vendor escalation refreshed when evidence changes or the case stalls.
The exact cadence varies. The anti-pattern is waiting for perfect knowledge. Silence invites customers to discover impact through their own failures and executives to obtain status by interrupting operators.
7. Automation, AI assistants, and the command boundary
Automation can improve incident response when it narrows repetitive work, preserves evidence, and executes bounded actions quickly. It can harm response when it acts faster than the organisation can understand or stop.
Knight Capital is the classic warning from outside reliability engineering. The SEC found that faulty automated trading code sent millions of erroneous orders in the first forty-five minutes of trading and caused a loss of more than USD 460 million.[17] Its relevance here has little to do with trading. It shows the speed at which automated action can turn a control gap into an organisational crisis.
In incident response, automation should be assigned an action tier:
| Tier | Examples | Command rule |
|---|---|---|
| Observe | Collect logs, snapshot dashboards, summarise tickets, correlate deploys | Allowed with provenance |
| Suggest | Recommend rollback, identify likely blast radius, draft customer update | Human owner accepts or rejects |
| Bounded action | Restart a stateless worker, rotate an exposed token, block known malicious IP | Pre-approved policy and automatic evidence record |
| Material action | Regional failover, data restore, bulk customer change, external disclosure | Named approval and command record |
| Critical action | Destructive deletion, privilege expansion, irreversible data change | Dual control or exceptional executive authority |
AI assistants belong in this table. They can be useful during incidents because they can summarise long chats, group related alerts, draft updates, search runbooks, produce timelines, and compare observed behaviour with known failure modes. They can also invent confidence, obscure uncertainty, follow malicious or irrelevant instructions in untrusted content, leak sensitive data into summaries, or make a brittle process look smarter than it is.
NIST's Generative AI Profile treats documentation, monitoring, human oversight, and governance as parts of AI risk management.[19] CISA's secure-by-design guidance emphasises manufacturer ownership of customer security outcomes, transparency, accountability, and leadership responsibility.[20] Those ideas translate directly into incident tooling: an AI assistant should not become the hidden commander.
A safe incident AI assistant should therefore:
- cite the underlying evidence for every factual claim;
- distinguish observed facts from hypotheses;
- show what sources it could not access;
- preserve the prompt, context, model, and output used for material decisions;
- avoid ingesting unnecessary secrets or personal data;
- require human ownership for recommendations;
- never approve its own material action;
- be disabled or degraded without removing human access to the incident record.
The test is not whether the assistant sounds calm. The test is whether it improves the evidence, authority, ownership, and communication channels without hiding responsibility.
8. Recovery, learning, and metrics beyond MTTR
Mean time to resolution is attractive because it compresses an incident into one number. It is also easy to misuse.
MTTR can improve because responders fixed the system faster. It can also improve because incidents are closed before recovery is complete, because customer impact is undercounted, because a long tail of data repair is excluded, or because teams avoid declaring incidents until certainty is high.
Recovery extends well beyond the end of alerting. It includes returning the system, data, customers, operators, and organisation to a known safe state. Public data-loss incidents such as GitLab's 2017 database outage show why recovery must include the data state as well as service availability.[15]
A recovery checklist should ask:
- Are user-facing symptoms back within SLO?
- Is data integrity verified?
- Are backlogs drained?
- Are workarounds still active?
- Were temporary permissions, firewall rules, feature flags, or capacity changes reverted?
- Are delayed jobs, billing effects, notifications, and customer-visible records correct?
- Did support receive closure language?
- Did customers receive a final update or incident report when appropriate?
- Is evidence preserved for post-incident review?
- Are follow-up actions owned and prioritised?
Google's SLO chapter argues for carefully defined service level indicators and objectives that measure aspects of service users care about, such as latency, error rate, throughput, and availability.[23] DORA's metrics combine delivery throughput and instability, including failed deployment recovery time, change fail rate, and deployment rework rate.[18] These approaches suggest a broader measurement set for incident command.
Interface metrics
The organisational interface can be measured directly.
| Metric | What it shows |
|---|---|
| Time to declare | Whether weak signals enter command quickly |
| Time to first customer-impact estimate | Whether evidence flows from telemetry and support to command |
| Time to first mitigation | Whether diagnosis is blocking harm reduction |
| Decision latency | How long material decisions wait for authority |
| Workstream ownership coverage | Percentage of active threads with named owner, next action, and check-in time |
| Communication freshness | Age of latest internal, executive, support, customer, and vendor update |
| Evidence provenance rate | Percentage of key claims linked to logs, dashboards, tickets, deploys, or vendor notices |
| Action collision count | Number of conflicting or uncoordinated system changes during incident |
| Handoff completeness | Whether incoming command can answer state, decisions, owners, risks, and clocks |
| Recovery debt | Temporary changes, backlogs, data repairs, and customer commitments left after closure |
| Recurrence with same contributing factors | Whether post-incident learning changed the system |
| Responder load | Hours awake, pages per person, context switches, and time to relief |
| Customer-visible impact | Error minutes, failed transactions, delayed work, data loss, or missed obligations |
These measures should not become a punishment dashboard. If a team delays declaration because the metric is used to shame them, the metric has damaged recovery. Use measures to improve the interface, not to rank heroes.
Post-incident learning
Google's postmortem guidance describes a postmortem as a written record of impact, actions, root causes, and follow-up actions, and emphasises blameless learning, broad review, and shared knowledge.[9] Richard Cook's classic human-factors paper argues that complex systems are hazardous, heavily defended, and dependent on practitioners' adaptive work; catastrophe usually emerges from multiple factors rather than a single isolated cause.[21]
Those ideas push incident reviews away from the comforting sentence, "Human error caused the incident." A better review asks:
- Why did the action make sense at the time?
- What information was missing, misleading, or inaccessible?
- Which defences failed or were bypassed?
- Which dependencies surprised us?
- Which command decisions helped recovery?
- Which communication gaps created extra work?
- Which controls made recovery slower but may be justified for security?
- Which action items will change the system rather than merely remind people to be careful?
The Facebook 2021 outage is useful because its public account notes a tradeoff: security hardening slowed recovery from an internally caused outage, but the company described the tradeoff as worth it and committed to stronger testing, drills, and resilience.[11] Good learning does not always produce a simple rule to remove friction. Sometimes it clarifies why friction exists and how to operate within it.
9. The strongest counterargument
The strongest counterargument is that incident command can become performance theatre.
A failing service needs engineers to fix it. Too much process can create meetings, forms, role labels, and approval queues that consume attention while users wait. A senior engineer may know the exact rollback needed, but a rigid command process may force them to wait for an Incident Commander who is less informed. A Communications Lead may polish language while the mitigation is obvious. A template may be filled with stale or speculative text because the process demands updates even when the facts have not changed. If every small issue becomes an incident, the organisation trains people to ignore incident rituals.
This counterargument is serious.
Incident command fails when it becomes more important than the incident. Common failure modes include:
- declaring too late because the process feels heavy;
- declaring too often until severity loses meaning;
- assigning roles without granting authority;
- treating the commander as a manager rather than a co-ordinator;
- excluding the person with local knowledge because they lack title;
- forcing all communication through one bottleneck;
- demanding root cause before mitigation;
- creating approval gates for reversible actions;
- turning the incident document into a reporting burden rather than a shared state tool;
- measuring only MTTR and thereby encouraging premature closure;
- using postmortems to assign blame in blameless language;
- allowing executives to use command channels for pressure rather than decisions.
The answer is proportional command.
A useful rule is: add structure at the point where it reduces coordination cost. If two people can solve a minor alert with no customer impact, a lightweight log and owner may be enough. If customer impact is growing, multiple teams are acting, external communication is needed, or authority is unclear, declare the incident and assign roles.
Process should be reversible. The Incident Commander should downshift, combine roles, close dormant channels, and release people. A high-quality command interface makes work easier for responders. If it regularly makes work harder, it should be treated as a process incident.
10. What this paper does not claim
This paper does not claim that incident command replaces technical expertise. The Operations Lead and domain experts remain essential.
It does not claim that every incident needs a large call, formal war room, or four named leads. Scale the interface to the risk.
It does not claim that diagnosis is unimportant. Understanding cause is necessary for safe recovery, prevention, and trust. The claim is that harm reduction often cannot wait for complete diagnosis.
It does not claim that all public incidents cited here failed because of poor incident command. The public sources do not support that broad judgement. They are used only for the specific claims described: dependencies, communication, recovery complexity, automation risk, and organisational coordination.
It does not claim that AI assistants are unsafe by definition. They can help incident teams. They should not become unaccountable decision makers.
It does not claim that MTTR is useless. It claims that MTTR alone is too narrow to describe recovery, learning, customer impact, or command quality.
It does not claim that process should dominate local judgement. The command interface exists to move useful judgement to the right place, not to suppress it.
Conclusion
Incident response is usually discussed in the language of speed: detect faster, mitigate faster, recover faster. Speed matters. Customers experience delay as failure.
But speed is an outcome of organisational design.
A team recovers faster when evidence reaches the people who can interpret it, when authority is known before the decision is needed, when every workstream has an owner, when customers and executives receive accurate updates without interrupting operators, when handoffs preserve context, when fatigue is managed, when automation is bounded, and when post-incident learning changes the system.
That is what incident command provides when it works.
It is not a heroic debugging ritual. It is an interface.
The interface has inputs: alerts, telemetry, tickets, reports, constraints, and uncertainty.
It has outputs: mitigations, decisions, communication, recovery tasks, evidence records, and learning.
It has failure modes: stale evidence, unclear authority, orphaned work, noisy communication, uncontrolled automation, exhausted responders, premature closure, and learning that never becomes change.
It can therefore be designed and measured.
The practical test is simple. During the next serious incident, can the organisation answer these questions without interrupting the people doing the mitigation?
- What do we know?
- What are we doing now?
- Who owns each thread?
- Who can approve the next material action?
- Who has been told what, and when is the next update?
- What remains unsafe after the alert clears?
If the answer is yes, incident command is doing its job. If the answer is no, the organisation does not merely have an incident. It has an interface failure.
About the author
Jason Doyle writes about reliable software, observability, applied AI, and practical controls for systems that influence human decisions. He publishes at jasondoyle.ie and can be contacted at contact@jasondoyle.ie.
References
- National Institute of Standards and Technology, Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Revision 3, April 2025, https://doi.org/10.6028/NIST.SP.800-61r3.
- National Institute of Standards and Technology, The NIST Cybersecurity Framework 2.0, 2024, https://doi.org/10.6028/NIST.CSWP.29.
- Cybersecurity and Infrastructure Security Agency, Federal Government Cybersecurity Incident and Vulnerability Response Playbooks, November 2021, https://www.cisa.gov/resources-tools/resources/federal-government-cybersecurity-incident-and-vulnerability-response-playbooks.
- Cybersecurity and Infrastructure Security Agency, Incident Reporting System, https://www.cisa.gov/resources-tools/resources/incident-reporting-system.
- Federal Emergency Management Agency, ICS Principle: Manageable Span of Control, https://emilms.fema.gov/is_0362a/groups/103.html.
- Google Site Reliability Engineering, Incident Management Guide, https://sre.google/resources/practices-and-processes/incident-management-guide/.
- Andrew Stribblehill, Managing Incidents, in Site Reliability Engineering, Google, https://sre.google/sre-book/managing-incidents/.
- Corey Adam Baye, Emergency Response, in Site Reliability Engineering, Google, https://sre.google/sre-book/emergency-response/.
- John Lunney and Sue Lueder, Postmortem Culture: Learning from Failure, in Site Reliability Engineering, Google, https://sre.google/sre-book/postmortem-culture/.
- Amazon Web Services, Summary of the Amazon S3 Service Disruption in the Northern Virginia (US-EAST-1) Region, 28 February 2017, https://aws.amazon.com/message/41926/.
- Santosh Janardhan, Meta Engineering, More details about the October 4 outage, 5 October 2021, https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/.
- Atlassian, April 2022 outage update, 18 April 2022, https://www.atlassian.com/blog/how-we-build/april-2022-outage-update.
- CrowdStrike, Falcon Content Update Remediation and Guidance Hub, 2024, https://www.crowdstrike.com/falcon-content-update-remediation-and-guidance-hub/.
- David Weston, Microsoft, Helping our customers through the CrowdStrike outage, 20 July 2024, https://blogs.microsoft.com/blog/2024/07/20/helping-our-customers-through-the-crowdstrike-outage/.
- GitLab, GitLab.com database incident, 1 February 2017, https://about.gitlab.com/blog/gitlab-dot-com-database-incident/.
- Cyber Safety Review Board, Review of the December 2021 Log4j Event, 2022, https://www.cisa.gov/resources-tools/resources/csrb-review-december-2021-log4j-event.
- United States Securities and Exchange Commission, SEC Charges Knight Capital With Violations of Market Access Rule, 16 October 2013, https://www.sec.gov/newsroom/press-releases/2013-222.
- DORA, DORA's software delivery metrics: the four keys, https://dora.dev/guides/dora-metrics/.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024, https://doi.org/10.6028/NIST.AI.600-1.
- Cybersecurity and Infrastructure Security Agency, Shifting the Balance of Cybersecurity Risk: Principles and Approaches for Secure by Design Software, 2023 and updated guidance, https://www.cisa.gov/resources-tools/resources/secure-by-design.
- Richard I. Cook, How Complex Systems Fail, 1998, https://how.complexsystems.fail/.
- United States Securities and Exchange Commission, Cybersecurity Risk Management, Strategy, Governance, and Incident Disclosure, Federal Register, 4 August 2023, https://www.federalregister.gov/documents/2023/08/04/2023-16194/cybersecurity-risk-management-strategy-governance-and-incident-disclosure.
- Google Site Reliability Engineering, Service Level Objectives, https://sre.google/sre-book/service-level-objectives/.