Start here
A good place to begin.
I write about operating cloud infrastructure, building teams, and using AI without giving up human judgment.
If this is your first visit, these four pieces are the clearest introduction to how I think and what I have worked on.
Scale
How we shipped 15 Tbps for OpenAI in 90 days
A first-person account of the architecture decisions, capacity work, global coordination, and GPU-based validation behind a high-stakes delivery.
Operations
Rewiring incident response, with AI in the loop
What changes when a growing cloud company replaces scattered incident habits with a system people can actually run.
Point of view
ITIL vs SRE: why the big clouds went their own way
A practical explanation of why hyperscalers operate differently, and which operational habits still matter.
Foundations
Service Levels: SLI, SLO, SLA
A compact guide to the reliability language customers and engineering teams need to share.
Then explore