Experience

Principal Site Reliability Engineer 2021 – present
Vela Technologies · San Francisco
  • Led migration of 200+ microservices from self-managed DC to AWS EKS, reducing P50 deployment time by 73% and saving $1.2M/yr in infrastructure.
  • Designed and implemented unified observability stack (OpenTelemetry + Prometheus + Grafana) covering 98% of services, enabling SLO-based alerting and reducing MTTR from 45min to 12min.
  • Built internal runbook automation platform (Go + Temporal) handling 60% of common incidents without human intervention.
  • Mentored 6 SREs, established incident command training and quarterly chaos days.
Senior Infrastructure Engineer 2018 – 2021
Stitch Labs (acquired) · Remote
  • Designed multi-region Kubernetes federation spanning us-east-1 / eu-west-1, achieving 99.99% uptime for inventory platform serving 10k+ merchants.
  • Reduced cloud spend by 34% ($800k/yr) via right-sizing, spot instances, and commitment-based purchasing; built FinOps dashboard with custom cost attribution.
  • Implemented canary deployments and渐进式交付 using Flagger + Istio, cutting deployment failure rate by 80%.
  • Wrote extensive Terraform modules for internal self-service (VPC, RDS, EKS clusters) adopted by 15 teams.
Systems Engineer / SRE 2015 – 2018
DigitalOcean · New York, NY
  • Part of core control-plane team: maintained provisioning and orchestration for fleet of 500k+ droplets, achieving 99.99% API availability.
  • Designed and built distributed health-check and auto-remediation system (Python + etcd) reducing manual page outs by 40%.
  • Led migration of monitoring from Nagios to Prometheus, establishing service-level dashboards and alerting with proper error budgets.
Junior Systems Administrator 2013 – 2015
Cornell University IT · Ithaca, NY
  • Managed Linux server fleet (RHEL/CentOS) for research computing; automated patching with Ansible and reduced vulnerability exposure window by 80%.
  • Built internal monitoring tool for HPC cluster utilization (Python + Graphite) adopted by 3 departments.

Open source & projects

sloctl · CLI for SLO management
Go · OpenMetrics · Prometheus · Kubernetes
Declarative SLO definition & validation tool. Used internally at Vela and adopted by 2 external organizations. ~1.2k GitHub stars.
chaos-patterns · collection of chaos experiments
Python · Litmus · AWS FIS · Terraform
Open-source library of reusable chaos scenarios (network latency, pod failure, AZ outage) with post-experiment analysis. Presented at SREcon 2023.
otel-collector-custom-processor
Go · OpenTelemetry Collector · Prometheus
Custom OpenTelemetry processor for tail-based sampling with dynamic priority based on service criticality. Deployed in production sampling >2TB traces/day.

Recognition

⚡ SREcon 2023 · best lightning talk
"Toil is a tax on velocity" – framework for measuring and reducing operational overhead.
🏆 Vela Engineering Impact Award (2022)
For leading infrastructure migration that improved platform reliability and reduced costs.
📝 USENIX ;login: author (2021)
Co-authored "Observability for the Rest of Us: Practical SLOs in a Microservice World".

Education

B.S. Computer Science · minor in Mathematics
Cornell University · College of Engineering
2009 – 2013 · cum laude (GPA 3.7)
Thesis: "Reliability in Distributed Storage: A Quantitative Analysis of Erasure Coding vs. Replication"