Resume

Site Reliability Engineer · Delray Beach, FL
[email protected] · GitHub · LinkedIn
Download PDF

Site Reliability Engineer focused on improving service reliability and performance through infrastructure as code, robust observability, and operational excellence. Skilled in designing and owning AWS infrastructure with Terraform and Terragrunt, from serverless platforms to production data pipelines. Experienced in defining SLIs/SLOs, reducing alert fatigue, and instrumenting applications to make root cause analysis routine. Trusted partner in cross-functional teams, bridging support, development, and platform engineering to build reliable, maintainable systems at scale.

Experience

Site Reliability Engineer 2 — ModMed

Apr 2023 – Present

Boca Raton, FL

  • Owned end-to-end infrastructure for a HIPAA-scope serverless clinical-transcription platform, authoring a ~75-resource Terraform module spanning private API Gateway, containerized Lambdas, and KMS-encrypted storage, with GitHub Actions-backed CI/CD and Datadog observability across multiple environments.
  • Designed an automated AWS data-extraction platform provisioning per-client S3 buckets, DataSync transfers, and MySQL exports from a single Terragrunt apply, collapsing a manual multi-step process into one command.
  • Operated and maintained a production Debezium/Kafka Connect CDC pipeline (SQL Server, AWS MSK, Schema Registry, DynamoDB, Databricks) delivering real-time healthcare data sync for a multi-tenant FHIR platform in a HIPAA-regulated environment.
  • Led the organization’s Datadog observability strategy, centralizing logging, building real-time service-health dashboards for engineering and support teams, and running sampling pipelines to reduce noise and ingestion cost while preserving signal.
  • Championed Datadog APM adoption beyond the original observability scope, instrumenting legacy applications so development teams could troubleshoot production issues directly, transforming root cause analysis from rare, inconclusive efforts into a standard, confident practice.
  • Developed custom Python-based Datadog checks for application and infrastructure components not covered by native integrations.
  • Standardized configuration management across legacy infrastructure with Ansible and reusable EC2 Image Builder pipelines, reducing drift between manually maintained servers.
  • Spearheaded automation initiatives by identifying repetitive tasks and implementing self-service forms, scripts, and workflow integrations, reducing team toil.
  • Defined and implemented SLIs/SLOs for core services, enabling proactive alerting and reliability tracking.
  • Reduced alert noise 80% by retiring non-actionable monitors and tuning noisy ones, curbing on-call fatigue.
  • Authored runbooks for every production alert, enabling any on-call engineer to triage without escalating to SMEs.
  • Rebuilt the on-call process in PagerDuty with clear escalation policies and service ownership.
  • Served as primary incident responder, leading triage, coordinating cross-functional response, and driving blameless postmortems.
  • Built CI/CD pipelines with Jenkins, GitHub Actions, and ArgoCD, implementing GitOps deployments to EKS (Kubernetes on AWS) with automated Docker builds pushed to AWS ECR.
  • Mentored engineering and support teams on observability best practices, driving org-wide Datadog adoption.

Support Services Engineer — ModMed

Aug 2021 – Apr 2023

Boca Raton, FL

  • Administered Windows infrastructure across AWS and on-prem datacenters, supporting nationwide healthcare operations.
  • Partnered with Cloud and Implementation teams to migrate legacy systems to AWS, enabling full datacenter decommissioning.
  • Maintained Python automation to schedule maintenance windows and suppress monitoring alerts via APIs (Google Calendar, LogicMonitor, Site24x7).
  • Built internal Google Apps Script tools to monitor enterprise-customer performance, improving operational visibility.
  • Developed PowerShell scripts to automate backup cleanups, reducing manual work and storage cost.
  • Collaborated with clients and vendors to implement and maintain VPN tunnels and SFTP integrations for secure data transfer.
  • Served as a technical liaison between Support and Cloud teams, enhancing incident response and service quality.
  • Mentored junior staff on IT fundamentals and systems architecture, contributing to promotions into cloud roles.

System Administrator — GA Telesis

Jul 2018 – Jul 2021

Fort Lauderdale, FL

  • Managed IT infrastructure across 8 global offices and 2 datacenters, ensuring high availability and reliability.
  • Led the migration from Exchange to Office 365, including training and rollout of Teams and OneDrive.
  • Enhanced disaster recovery by documenting restore playbooks and implementing automated backup testing.
  • Led team through rapid transformation to allow employees to work remotely.
  • Strengthened security by rolling out MFA, security-awareness training, and endpoint protection upgrades.
  • Optimized onboarding with standardized communication and documentation, improving new hire experience.
  • Deployed AWS S3 for marketing media and backup repositories.

IT Specialist — ExamSoft

May 2017 – Jun 2018

Delray Beach, FL

  • Supported corporate infrastructure and resolved internal IT issues via Jira Helpdesk.
  • Scripted automations to streamline onboarding and logging for Engineering.
  • Built a remote-access VPN enabling engineers to reach internal infrastructure while working remotely.
  • Maintained IT documentation and provided remote technical support.

IT Administrator — Sandy James Catering

Aug 2012 – May 2017

West Palm Beach, FL

  • Built and maintained custom software — including a Flask-based phone system — and administered the company’s ERP system.
  • Managed IT infrastructure and user support for office and remote staff.

Education

  • B.S., Applied Computer Science — Troy University (2018)

Skills

Cloud & Infrastructure
AWS (EC2, S3, Route 53, IAM, EKS, ECS, Lambda, API Gateway, DataSync, EC2 Image Builder), VMware vSphere, Proxmox VE, ESXi, Nutanix
Containers & Orchestration
Docker, Kubernetes, Helm, Kustomize
Infrastructure as Code & Automation
Terraform, Terragrunt, Ansible, AWX
CI/CD & GitOps
Jenkins, GitHub Actions, ArgoCD, GitOps
Observability & Monitoring
Datadog (APM, Logs, Dashboards, Custom Checks, SLOs), Prometheus, Grafana, LogicMonitor, Site24x7
Scripting & Languages
Python, PowerShell, Bash, JavaScript, SQL, Google Apps Script
Databases & Streaming
Microsoft SQL Server, MySQL, DynamoDB, Amazon DocumentDB, Kafka (AWS MSK, Kafka Connect, Debezium)
Secrets Management
AWS Secrets Manager, SSM Parameter Store
Version Control
Git, GitHub, Bitbucket
Operating Systems
Linux (Ubuntu, CentOS, Amazon Linux), Windows Server, macOS
Security & Endpoint Protection
CrowdStrike, Sophos, Cisco AMP, Qualys, TAEGIS (Secureworks), GuardiCore
Networking & Remote Access
VPN, SFTP, Cisco Meraki, Brocade
Backup & Recovery
Veeam, Commvault, N2WS
ITSM & Collaboration
Jira, Confluence, PagerDuty, Runbooks, Incident Management
System Services
Active Directory, WSUS