Writing


Writing.

Essays, notes, and longer posts — 16 pieces so far.

AI data centers: the five hard problems money cannot buy away

The AI buildout is not mainly a real-estate problem. It is a race to integrate power, cooling, equipment, permits, and useful compute at the same time.

ai
Keeping GPUs Fed

The GPU does not care that object storage is durable and scalable. It cares whether the next batch of data arrives before it goes idle.

ai
Should senior leadership attend a PIR?

For major incidents, a Post Incident Review is not an operations meeting. It is where leaders remove the blockers that keep the service from becoming safer.

incident-management
Why GitHub feels less reliable lately

GitHub is not having one outage problem. Its recent incident reports show the difficult middle of a platform transformation.

sre
Your TLS rotation is not reliable until production proves it

Automation renews a certificate. Observability proves every endpoint is serving it and customers can complete a TLS handshake.

sre
Cloud provider postmortems: volume vs depth

GCP publishes 100+ postmortems a year. AWS publishes almost none. Azure has become the transparency leader. What each posture reveals about engineering culture, and what SREs should steal from all three.

sre
When Atlas meets the hyperscale

Atlassian recommends consolidation. Hyperscalers use many. Both are right for different problems. Five real reasons to split, and what works at each scale.

sre
Direction, then review: my pattern for using AI at work

How I use Claude and ChatGPT like junior developers on my team. Three concrete workflows from this month, and a five-point checklist for reviewing AI output.

ai
ITIL vs SRE: why the big clouds went their own way

The big clouds don't run ITIL. Five assumptions ITIL makes that break at hyperscaler scale, and what AWS, Azure, GCP, and OCI use instead.

sre
Rewiring incident response, with AI in the loop

Six weeks in at a new company, rewiring incident response from one Slack thread to a working stack, with AI compressing the parts that used to take a quarter.

sre
How we shipped 15 Tbps for OpenAI in 90 days (Session 1 of 3)

OpenAI wanted a second Object Storage instance, in customer-facing production, at 15 Tbps, in three months. Session 1 covers the first week: closing the architecture.

sre
Storage at scale: what I actually watched

For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.

sre
Five rules for running an incident

The difference between a 10-minute incident and a 3-hour outage is rarely technical. Five things I wish every on-call team locked in before their first big page.

sre
Three 5xx and one 4xx: the codes I actually care about

How I read 500, 502, 503, and 429 in production at scale, and what each one is really telling you.

sre
Operational Readiness: The Review That Catches Problems

A short, verifiable checklist for production launches. What to ask, why ORRs become theater, and where AI helps.

sre
Service Levels: SLI, SLO, SLA

What SLI, SLO, and SLA actually mean, why the order matters, and where AI helps.

sre