Writing
Writing.
Essays, notes, and longer posts — 16 pieces so far.
Also: daily learning notes · RSS · Start here
The AI buildout is not mainly a real-estate problem. It is a race to integrate power, cooling, equipment, permits, and useful compute at the same time.
The GPU does not care that object storage is durable and scalable. It cares whether the next batch of data arrives before it goes idle.
For major incidents, a Post Incident Review is not an operations meeting. It is where leaders remove the blockers that keep the service from becoming safer.
GitHub is not having one outage problem. Its recent incident reports show the difficult middle of a platform transformation.
Automation renews a certificate. Observability proves every endpoint is serving it and customers can complete a TLS handshake.
GCP publishes 100+ postmortems a year. AWS publishes almost none. Azure has become the transparency leader. What each posture reveals about engineering culture, and what SREs should steal from all three.
Atlassian recommends consolidation. Hyperscalers use many. Both are right for different problems. Five real reasons to split, and what works at each scale.
How I use Claude and ChatGPT like junior developers on my team. Three concrete workflows from this month, and a five-point checklist for reviewing AI output.
The big clouds don't run ITIL. Five assumptions ITIL makes that break at hyperscaler scale, and what AWS, Azure, GCP, and OCI use instead.
Six weeks in at a new company, rewiring incident response from one Slack thread to a working stack, with AI compressing the parts that used to take a quarter.
OpenAI wanted a second Object Storage instance, in customer-facing production, at 15 Tbps, in three months. Session 1 covers the first week: closing the architecture.
For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.
The difference between a 10-minute incident and a 3-hour outage is rarely technical. Five things I wish every on-call team locked in before their first big page.
How I read 500, 502, 503, and 429 in production at scale, and what each one is really telling you.
A short, verifiable checklist for production launches. What to ask, why ORRs become theater, and where AI helps.
What SLI, SLO, and SLA actually mean, why the order matters, and where AI helps.