REALITYWIPE

RESEARCH_FILE

The Science of Engineering Incidents and Identity

"Who broke prod?"

SEE THE PRACTICE

Turn a thought this research explains into one clear move.

THE THOUGHT

I inherited this mess and now everyone thinks it's my fault

YOUR RECORDED RESPONSE

This bug has a birthdate, and it's before my first commit here.

ONE PRIVATE MOVE

Pick one open ticket on inherited code and write a two-line note: what the bug is, and roughly when the code was introduced. Don't send it anywhere yet.

feels like the right question after an outage. Systems research gives a less satisfying but far more useful answer: complex systems fail through the alignment of many small conditions, not through one reckless actor (Reason 1990; Google SRE 2016). Meanwhile the people most responsible for keeping those systems alive — the ones doing infrastructure work, on-call rotations, and glue work — are systematically under-credited (Champion et al. 2024). And now a new identity threat layers on top: AI tools that promise to replace engineers are trusted by a shrinking minority of developers (Stack Overflow 2025). The old script — 'someone here is the villain' — is almost always the wrong read.

~75%of Stack Overflow 2025 survey respondents report distrust or low trust in AI-generated code — the lowest level of developer trust in AI output ever recorded in the survey5+ layersof defensive barriers that must all have simultaneous holes for a complex system failure to occur, per Reason's Swiss cheese model — not one bad actor, not one missed checkSystematicunder-recognition of infrastructure, documentation, and review work in open source ecosystems — Champion et al. 2024 find these contributions are structurally invisible compared to feature commits in most recognition systems1 chapterin Google's SRE Book dedicated entirely to postmortem culture — codifying blamelessness as an organizational requirement, not a nice-to-have, because blame actively degrades the information quality needed to prevent future incidents

How the science changed

  1. 1983

    Arlie Hochschild's The Managed Heart coins the concept of emotional labor — work that requires managing one's feelings as part of the job but goes unmeasured in any formal accounting. The concept later anchors research on invisible glue work in software teams.

  2. 1990

    James Reason publishes Human Error, introducing the Swiss cheese model of accident causation: failures happen when holes in multiple defensive layers align, not because one person made one mistake. The model becomes foundational to aviation, medicine, and later software reliability.

  3. 2006

    Sidney Dekker's The Field Guide to Understanding Human Error reframes 'human error' as a symptom, not a cause — the engineer who clicked the wrong button did so inside a system that allowed or encouraged that action. Blame, Dekker argues, stops investigations at the wrong layer.

  4. 2012

    John Allspaw and Jesse Robbins popularize blameless postmortems at Etsy, arguing that engineers who fear punishment hide information — and hidden information kills reliability. Their approach later influences Google's Site Reliability Engineering practice.

  5. 2016

    Google publishes its SRE Book, dedicating a full chapter to postmortem culture. The chapter codifies blamelessness as an organizational policy: 'the primary goals of writing a postmortem are to ensure the incident is documented, all contributing root causes are understood, and effective preventive actions are put in place.'

  6. 2018

    Forsgren, Humble, and Kim publish Accelerate, correlating organizational practices with software delivery performance across thousands of teams. Their data shows that high-performing teams are distinguished not by blame-free heroes but by psychological safety, fast feedback loops, and distributed ownership of reliability.

  7. 2024

    Champion et al. publish a large-scale empirical study of invisible labor in open source software ecosystems: infrastructure maintenance, documentation, code review, and community management are systematically under-recognized compared to feature commits — and the gap disadvantages engineers who specialize in reliability over novelty.

What people believe vs. what the data shows

The beliefWhen a production incident happens, there is always one root cause and one person responsible for it.

The dataReason's Swiss cheese model and Dekker's field guide both show that complex system failures require multiple holes in multiple defensive layers to line up simultaneously. Google's SRE Book explicitly rejects single-root-cause thinking: incidents are the product of latent conditions, not individual villains.

The beliefIf you want accountability after an incident, someone needs to be blamed and punished.

The dataAllspaw and Robbins documented the opposite: fear of punishment causes engineers to withhold information, which makes future incidents more likely. Google's blameless postmortem framework shows that real accountability — documented causes, system fixes, preventive actions — is actually incompatible with individual blame.

The beliefEngineers who work on infrastructure, on-call rotations, and documentation are doing less valuable work than those shipping visible features.

The dataChampion et al.'s 2024 empirical study of open source ecosystems found that infrastructure maintenance, documentation, and code review are systematically under-credited even though they are the foundation that visible features run on. Forsgren et al.'s Accelerate data shows that reliability work — not feature velocity alone — predicts high organizational performance.

The beliefThe right response to an incident is to find the person who was on-call and hold them responsible.

The dataGoogle's SRE postmortem framework explicitly separates the person responding from the conditions enabling the failure. The on-call engineer is often the one who surfaced the problem, not the one who created the conditions. Treating them as the cause discourages future responders from engaging honestly.

The beliefAI coding tools are now trusted enough that engineers who feel anxious about AI replacing them are being irrational.

The dataThe 2025 Stack Overflow Developer Survey found that developer trust in AI output is at an all-time low — the majority of developers do not fully trust AI-generated code. Anxiety about AI tools is a well-grounded professional response to an immature technology being marketed as mature.

TEST_YOURSELF · How well do you know this science?

  1. 01 What does James Reason's Swiss cheese model say about how complex systems fail?

    Reason's model shows complex system failures are multi-causal: a series of holes in otherwise-robust defensive layers must line up for an accident to penetrate to the outcome. No single person creates all the holes. source

  2. 02 According to Google's SRE Book chapter on postmortem culture, what is the primary purpose of a blameless postmortem?

    Google's SRE Book states that the primary goals of a postmortem are to ensure the incident is documented, all contributing root causes are understood, and effective preventive actions are put in place — not to assign fault to individuals. source

  3. 03 What did Champion et al.'s 2024 study find about infrastructure and glue work in open source software ecosystems?

    Champion et al. found that infrastructure maintenance, documentation, code review, and community management are structurally invisible in most recognition systems — despite being essential to the ecosystem's functioning — while feature commits attract disproportionate credit. source

  4. 04 According to the 2025 Stack Overflow Developer Survey, what is happening to developer trust in AI-generated code?

    The 2025 Stack Overflow Developer Survey found developer trust in AI output at an all-time low — the majority of developers do not fully trust AI-generated code, despite the rapid proliferation of AI coding tools. source

  5. 05 What did Allspaw and Robbins find about the effect of blame culture on engineering incident investigations?

    Allspaw and Robbins documented that fear of punishment causes engineers to hide information about what happened during an incident. That hidden information prevents the organization from understanding and fixing the real conditions — making the next incident more likely. source

The researchers behind it

Named mechanisms

Scripts this research explains

Related old scripts

Related self-checks

Related science

Researchers to explore

Related mechanisms

by the numbersresearch timelinetest your knowledgeAll research