<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Sridhar Rajarao</title><description>Sridhar Rajarao&apos;s writing — notes, essays, and longer-form posts.</description><link>https://sridharrajarao.com/</link><language>en-us</language><item><title>AI data centers: the five hard problems money cannot buy away</title><link>https://sridharrajarao.com/blog/ai-datacenter-buildout-five-issues/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/ai-datacenter-buildout-five-issues/</guid><description>The AI buildout is not mainly a real-estate problem. It is a race to integrate power, cooling, equipment, permits, and useful compute at the same time.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Keeping GPUs Fed</title><link>https://sridharrajarao.com/blog/gpus-need-fast-object-storage/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/gpus-need-fast-object-storage/</guid><description>The GPU does not care that object storage is durable and scalable. It cares whether the next batch of data arrives before it goes idle.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Should senior leadership attend a PIR?</title><link>https://sridharrajarao.com/blog/should-senior-leadership-attend-pir/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/should-senior-leadership-attend-pir/</guid><description>For major incidents, a Post Incident Review is not an operations meeting. It is where leaders remove the blockers that keep the service from becoming safer.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Why GitHub feels less reliable lately</title><link>https://sridharrajarao.com/blog/why-github-feels-less-reliable/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/why-github-feels-less-reliable/</guid><description>GitHub is not having one outage problem. Its recent incident reports show the difficult middle of a platform transformation.</description><pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Your TLS rotation is not reliable until production proves it</title><link>https://sridharrajarao.com/blog/tls-rotation-observability/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/tls-rotation-observability/</guid><description>Automation renews a certificate. Observability proves every endpoint is serving it and customers can complete a TLS handshake.</description><pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Cloud provider postmortems: volume vs depth</title><link>https://sridharrajarao.com/blog/cloud-postmortems-volume-vs-depth/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/cloud-postmortems-volume-vs-depth/</guid><description>GCP publishes 100+ postmortems a year. AWS publishes almost none. Azure has become the transparency leader. What each posture reveals about engineering culture, and what SREs should steal from all three.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate></item><item><title>When Atlas meets the hyperscale</title><link>https://sridharrajarao.com/blog/jsm-projects-atlassian-vs-hyperscalers/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/jsm-projects-atlassian-vs-hyperscalers/</guid><description>Atlassian recommends consolidation. Hyperscalers use many. Both are right for different problems. Five real reasons to split, and what works at each scale.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Direction, then review: my pattern for using AI at work</title><link>https://sridharrajarao.com/blog/direction-then-review/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/direction-then-review/</guid><description>How I use Claude and ChatGPT like junior developers on my team. Three concrete workflows from this month, and a five-point checklist for reviewing AI output.</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate></item><item><title>ITIL vs SRE: why the big clouds went their own way</title><link>https://sridharrajarao.com/blog/itil-vs-sre/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/itil-vs-sre/</guid><description>The big clouds don&apos;t run ITIL. Five assumptions ITIL makes that break at hyperscaler scale, and what AWS, Azure, GCP, and OCI use instead.</description><pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Rewiring incident response, with AI in the loop</title><link>https://sridharrajarao.com/blog/rewiring-incident-response/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/rewiring-incident-response/</guid><description>Six weeks in at a new company, rewiring incident response from one Slack thread to a working stack, with AI compressing the parts that used to take a quarter.</description><pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate></item><item><title>How we shipped 15 Tbps for OpenAI in 90 days (Session 1 of 3)</title><link>https://sridharrajarao.com/blog/openai-15-tbps-session-1/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/openai-15-tbps-session-1/</guid><description>OpenAI wanted a second Object Storage instance, in customer-facing production, at 15 Tbps, in three months. Session 1 covers the first week: closing the architecture.</description><pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Storage at scale: what I actually watched</title><link>https://sridharrajarao.com/blog/storage-at-scale/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/storage-at-scale/</guid><description>For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.</description><pubDate>Thu, 28 May 2026 00:00:00 GMT</pubDate></item><item><title>Five rules for running an incident</title><link>https://sridharrajarao.com/blog/running-an-incident/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/running-an-incident/</guid><description>The difference between a 10-minute incident and a 3-hour outage is rarely technical. Five things I wish every on-call team locked in before their first big page.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate></item><item><title>Three 5xx and one 4xx: the codes I actually care about</title><link>https://sridharrajarao.com/blog/http-codes-at-scale/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/http-codes-at-scale/</guid><description>How I read 500, 502, 503, and 429 in production at scale, and what each one is really telling you.</description><pubDate>Tue, 26 May 2026 00:00:00 GMT</pubDate></item><item><title>Operational Readiness: The Review That Catches Problems</title><link>https://sridharrajarao.com/blog/operational-readiness/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/operational-readiness/</guid><description>A short, verifiable checklist for production launches. What to ask, why ORRs become theater, and where AI helps.</description><pubDate>Mon, 25 May 2026 00:00:00 GMT</pubDate></item><item><title>Service Levels: SLI, SLO, SLA</title><link>https://sridharrajarao.com/blog/service-levels/</link><guid isPermaLink="true">https://sridharrajarao.com/blog/service-levels/</guid><description>What SLI, SLO, and SLA actually mean, why the order matters, and where AI helps.</description><pubDate>Sun, 24 May 2026 00:00:00 GMT</pubDate></item></channel></rss>