Work
Work.
Current work, selected initiatives, writing, talks, and career.
Skip to
now,
featured work,
initiatives,
side projects,
publications,
speaking, or
career.
Current role
At Nscale.
Principal Infrastructure Engineer · Jun 2026 – Present Making GPU cloud operations easier to run at scale.
Streamlining incident and change management for GPU cloud operations.
Establishing on-call tooling, change templates, and the Grafana →
PagerDuty → Jira alert path. Also setting up Major Incident
guidelines so teams can respond with clarity when the stakes are high.
Featured work
15 Tbps in 90 days.
OpenAI · AI Infrastructure The first custom OCI Object Storage instance at production scale.
A $100M engagement with 15 Tbps of required bandwidth, delivered in
three months. The work involved a new network architecture, capacity
across the storage stack, and load validation from customer GPUs.
$100M · 15 Tbps · 3 months Selected initiatives
Programs I’m proud of.
Org Building The platform organization (5 → 45 engineers)
Grew SRE and Platform Engineering for OCI Object Storage from 5 to 45 across four geographies. Shifted the operating model so service teams consumed platforms instead of building their own.
5 → 45 engineers · 4 geographies Reliability Operating-model transformation
Production rollbacks dropped 60% per service. Build times dropped 80%. Change-caused incidents dropped 50%. The leverage compounds when service teams stop reinventing region builds, pre-prod qual, and fleet management.
60% / 80% / 50% reductions AIOps AI-driven SRE automation
Built AI-driven support and automation workflows across the incident lifecycle: detection, triage, remediation. Reduced ticket volume by 30% and improved response times by 50%.
30% fewer tickets · 50% faster responses Scale Region build framework
Built the repeatable region-build framework that enabled OCI expansion from 4 to 100+ regions, with consistent reliability and operational readiness at each launch.
4 → 100+ regions Compliance Fleet compliance and security
Automated patching and vulnerability management for the 50,000+ node fleet, achieving 100% compliance with FedRAMP standards without slowing engineering velocity.
50K+ nodes · 100% FedRAMP Side projects
Things I’m building.
iOS · Android · Web TrackMyGPU
Live cloud GPU pricing across 10 providers (AWS, GCP, Azure, Lambda,
RunPod, Vast.ai, CoreWeave, Paperspace, Hyperstack, Salad). A
Claude-powered advisor picks the cheapest reliable rig for your
workload, with push alerts when prices drop.
iOS · Android · Web Tab — Meals on Tab
A simple way to send a restaurant meal to a friend or pay one
forward for someone in the community. Diners can also request their
favorite restaurant and help bring it onto the platform.
Podcast · Video SRE(E) — the show
AI × SRE, every day, on the mic. Short daily takes on
outages, tools, and AI agents in production.
Publications
Longer-form writing.
Daily notes and shorter posts on the Writing page.
Career
Where I’ve been.
Jun 2026 – Present Principal Infrastructure Engineer · Nscale Streamlining incident and change management for GPU cloud operations. Establishing on-call tooling (PagerDuty, Jira, change templates), wiring the Grafana → PagerDuty → Jira alert path, and setting up Major Incident guidelines.
Current 2020 – Mar 2026 Director, Site Reliability Engineering · Oracle Cloud Infrastructure (OCI) Led the global SRE organization for OCI Object Storage — 45+ engineers and 3 managers across the US, India, UK, and Mexico, supporting 100+ regions and a 50,000+ node fleet. Owned end-to-end service reliability, incident lifecycle, automation strategy, and operational scalability for a business that grew to $600M ARR, 10 EB of customer data, and 100% YoY growth.
Past Mar 2014 – Mar 2020 SRE Manager · Oracle Inc, USA Led release and operational readiness for Object Storage across multi-region deployments. Built the region-build framework that enabled OCI expansion from 4 to 100+ regions, and developed automation for fleet patching across 30,000+ nodes.
Past 1993 – 2014 Enterprise consulting · Financial, banking, healthcare, security platforms Twenty-one years delivering enterprise implementations across regulated industries. Consistent on-time delivery with measurable defect reduction.
Past Download CV