Work


Work.

Current projects, past initiatives, writing, talks, and career. Skip to projects, initiatives, publications, speaking, or career.

Projects


Building now.

iOS · Android · Web

GPU Tracker

Live cloud GPU pricing across 10 providers (AWS, GCP, Azure, Lambda, RunPod, Vast.ai, CoreWeave, Paperspace, Hyperstack, Salad). A Claude-powered advisor picks the cheapest reliable rig for your workload, with push alerts when prices drop.

Podcast · Video

SRE(E) — the show

AI × SRE, every day, on the mic. Short daily takes on outages, tools, and AI agents in production.

Initiatives


Programs I’m proud of.

Org Building

The platform organization (5 → 45 engineers)

Grew SRE and Platform Engineering for OCI Object Storage from 5 to 45 across four geographies. Shifted the operating model so service teams consumed platforms instead of building their own.

5 → 45 engineers · 4 geographies
Reliability

Operating-model transformation

Production rollbacks dropped 60% per service. Build times dropped 80%. Change-caused incidents dropped 50%. The leverage compounds when service teams stop reinventing region builds, pre-prod qual, and fleet management.

60% / 80% / 50% reductions
AI Infrastructure

OpenAI: the first custom Object Storage at production scale

$100M engagement, 15 Tbps bandwidth, delivered in three months. Taught me what AI workloads actually do to storage at scale and shaped most of my thinking on where the field is going.

$100M · 15 Tbps · 3 months
AIOps

AI-driven SRE automation

Built AI-driven support and automation workflows across the incident lifecycle: detection, triage, remediation. Reduced ticket volume by 30% and improved response times by 50%.

30% fewer tickets · 50% faster responses
Scale

Region build framework

Built the repeatable region-build framework that enabled OCI expansion from 4 to 100+ regions, with consistent reliability and operational readiness at each launch.

4 → 100+ regions
Compliance

Fleet compliance and security

Automated patching and vulnerability management for the 50,000+ node fleet, achieving 100% compliance with FedRAMP standards without slowing engineering velocity.

50K+ nodes · 100% FedRAMP

Publications


Longer-form writing.

Daily notes and shorter posts on the Writing page.

Speaking


Talks.

Career


Where I’ve been.

Jun 2026 – Present
Principal Infrastructure Engineer · Nscale

Streamlining incident and change management for GPU cloud operations. Establishing on-call tooling (PagerDuty, Jira, change templates), wiring the Grafana → PagerDuty → Jira alert path, and setting up Major Incident guidelines.

Current
2020 – Mar 2026
Director, Site Reliability Engineering · Oracle Cloud Infrastructure (OCI)

Led the global SRE organization for OCI Object Storage — 45+ engineers and 3 managers across the US, India, UK, and Mexico, supporting 100+ regions and a 50,000+ node fleet. Owned end-to-end service reliability, incident lifecycle, automation strategy, and operational scalability for a business that grew to $600M ARR, 10 EB of customer data, and 100% YoY growth.

Past
Mar 2014 – Mar 2020
SRE Manager · Oracle Inc, USA

Led release and operational readiness for Object Storage across multi-region deployments. Built the region-build framework that enabled OCI expansion from 4 to 100+ regions, and developed automation for fleet patching across 30,000+ nodes.

Past
1993 – 2014
Enterprise consulting · Financial, banking, healthcare, security platforms

Twenty-one years delivering enterprise implementations across regulated industries. Consistent on-time delivery with measurable defect reduction.

Past

Download CV