REALITYWIPE

BY_THE_NUMBERS

The Science of Engineering Incidents and Identity: Why It's Rarely the Villain's Fault

"Who broke prod?" feels like the right question after an outage. Systems research gives a less satisfying but far more useful answer: complex systems fail through the alignment of many small conditions, not through one reckless actor (Reason 1990; Google SRE 2016). Meanwhile the people most responsible for keeping those systems alive — the ones doing infrastructure work, on-call rotations, and glue work — are systematically under-credited (Champion et al. 2024). And now a new identity threat layers on top: AI tools that promise to replace engineers are trusted by a shrinking minority of developers (Stack Overflow 2025). The old script — 'someone here is the villain' — is almost always the wrong read.

~75%of Stack Overflow 2025 survey respondents report distrust or low trust in AI-generated code — the lowest level of developer trust in AI output ever recorded in the survey5+ layersof defensive barriers that must all have simultaneous holes for a complex system failure to occur, per Reason's Swiss cheese model — not one bad actor, not one missed checkSystematicunder-recognition of infrastructure, documentation, and review work in open source ecosystems — Champion et al. 2024 find these contributions are structurally invisible compared to feature commits in most recognition systems1 chapterin Google's SRE Book dedicated entirely to postmortem culture — codifying blamelessness as an organizational requirement, not a nice-to-have, because blame actively degrades the information quality needed to prevent future incidents
View full research filetest your knowledgeAll research