·  sre, cloud, postmortems, incident-management


Cloud provider postmortems: volume vs depth

GCP publishes 100+ postmortems a year. AWS publishes almost none. Azure has become the transparency leader. What each posture reveals about engineering culture, and what SREs should steal from all three.

Every major cloud provider publishes some form of post-incident writeup. How much they publish, how fast, and how deep tells you a lot about their engineering culture. It varies more than you would expect.

Postmortem practice is a leading indicator of engineering discipline. What you publish is what you take seriously.

Google Cloud: high volume, fast, systematic

GCP publishes more than 100 postmortems a year. They post fast. They use the same template every time. On volume and speed, they lead.

The trade-off shows in depth. GCP’s postmortems can feel process-driven. Some have shifted between preliminary and final versions, with early architectural concerns quietly dropped in the later version. High cadence, medium average depth.

Azure: most transparent overall

Azure has become the leader on incident transparency and accountability over the last few years. Their post-incident writeups tend to be detailed, blameless in tone, and treat the customer as a partner. This is a recent posture. Azure of 2018 was not this.

If you are building your own postmortem template today, Azure’s 2025 and 2026 examples are the best public reference.

AWS: selective, low volume, occasionally legendary

AWS publishes Post-Event Summaries at aws.amazon.com/message/<id>/, only for issues with broad customer impact. They publish far fewer than the others. Individual customer impact goes through the Personal Health Dashboard and account managers, not public writeups.

The trade-off cuts the other way. The AWS 2017 S3 postmortem became a masterclass. It is still cited in SRE training material. The DynamoDB US-EAST-1 writeup from October 2025 is the same tradition. AWS publishes rarely, but when they do, the writing is deep enough to shape how the industry thinks about failure.

Volume is not depth

The obvious ranking (GCP most, AWS least) misses the point. 100 shallow postmortems teach the industry less than three legendary ones. AWS is playing a different game: fewer public writeups, higher average signal per writeup, more information flowing to paying customers privately.

None of this is neutral. Each posture reflects a strategic choice.

  • GCP: transparency as marketing. Every incident, publicly documented.
  • Azure: transparency as customer trust rebuild. Response to years of “Azure is opaque” complaints.
  • AWS: transparency as selective signal. Save public writeups for the moments that reshape the industry.

What SREs should actually steal

Look at what each provider does well:

  • From GCP: the cadence and the template. If you cannot publish quickly, you will never publish.
  • From Azure: the tone. Blameless writing, root-cause depth, and treating the customer as an audience worth respecting.
  • From AWS: the discipline. Not every incident deserves a public writeup. Save the depth for the ones that would actually teach the world something.

The takeaway

Postmortem practice reflects how a company thinks about failure. High volume signals cultural investment in documentation. High depth signals cultural investment in learning. High speed signals cultural investment in trust. No single provider gets all three right. If you are building your team’s postmortem practice, steal from all three.


Sources

← All writing