Projects
Building now.
iOS · Android · Web GPU Tracker
Live cloud GPU pricing across 10 providers (AWS, GCP, Azure, Lambda,
RunPod, Vast.ai, CoreWeave, Paperspace, Hyperstack, Salad). A
Claude-powered advisor picks the cheapest reliable rig for your
workload, with push alerts when prices drop.
Podcast · Video SRE(E) — the show
AI × SRE, every day, on the mic. Short daily takes on
outages, tools, and AI agents in production.
Initiatives
Programs I’m proud of.
Org Building The platform organization (5 → 45 engineers)
Grew SRE and Platform Engineering for OCI Object Storage from 5 to 45 across four geographies. Shifted the operating model so service teams consumed platforms instead of building their own.
5 → 45 engineers · 4 geographies Reliability Operating-model transformation
Production rollbacks dropped 60% per service. Build times dropped 80%. Change-caused incidents dropped 50%. The leverage compounds when service teams stop reinventing region builds, pre-prod qual, and fleet management.
60% / 80% / 50% reductions AI Infrastructure OpenAI: the first custom Object Storage at production scale
$100M engagement, 15 Tbps bandwidth, delivered in three months. Taught me what AI workloads actually do to storage at scale and shaped most of my thinking on where the field is going.
$100M · 15 Tbps · 3 months AIOps AI-driven SRE automation
Built AI-driven support and automation workflows across the incident lifecycle: detection, triage, remediation. Reduced ticket volume by 30% and improved response times by 50%.
30% fewer tickets · 50% faster responses Scale Region build framework
Built the repeatable region-build framework that enabled OCI expansion from 4 to 100+ regions, with consistent reliability and operational readiness at each launch.
4 → 100+ regions Compliance Fleet compliance and security
Automated patching and vulnerability management for the 50,000+ node fleet, achieving 100% compliance with FedRAMP standards without slowing engineering velocity.
50K+ nodes · 100% FedRAMP Publications
Longer-form writing.
Daily notes and shorter posts on the Writing page.
Career
Where I’ve been.
Jun 2026 – Present Principal Infrastructure Engineer · Nscale Streamlining incident and change management for GPU cloud operations. Establishing on-call tooling (PagerDuty, Jira, change templates), wiring the Grafana → PagerDuty → Jira alert path, and setting up Major Incident guidelines.
Current 2020 – Mar 2026 Director, Site Reliability Engineering · Oracle Cloud Infrastructure (OCI) Led the global SRE organization for OCI Object Storage — 45+ engineers and 3 managers across the US, India, UK, and Mexico, supporting 100+ regions and a 50,000+ node fleet. Owned end-to-end service reliability, incident lifecycle, automation strategy, and operational scalability for a business that grew to $600M ARR, 10 EB of customer data, and 100% YoY growth.
Past Mar 2014 – Mar 2020 SRE Manager · Oracle Inc, USA Led release and operational readiness for Object Storage across multi-region deployments. Built the region-build framework that enabled OCI expansion from 4 to 100+ regions, and developed automation for fleet patching across 30,000+ nodes.
Past 1993 – 2014 Enterprise consulting · Financial, banking, healthcare, security platforms Twenty-one years delivering enterprise implementations across regulated industries. Consistent on-time delivery with measurable defect reduction.
Past Download CV