Serve as an experienced individual contributor responsible for the availability, performance, and reliability of a large federal enterprise cloud platform that operates around the clock
Help meet scope, schedule, and delivery requirements while shaping the platform's reliability strategy
Defining and maintaining service level objectives (SLOs), service level indicators, and error budgets, and driving the platform toward them
Designing and operating observability across metrics, logging, tracing, and alerting
Leading incident response and on-call practices, including escalation, mitigation, and time-to-recovery improvements
Driving blameless postmortems and systemic reliability improvements
Engineering automation to eliminate toil and improve operational efficiency
Self-directed design of reliable cloud infrastructure (AWS) and Kubernetes (Amazon EKS)
Building reusable modules and mentoring engineers on reliability practices
Presenting design documents and system diagrams to stakeholders
Participating in technical depth interviews with new candidates
Requirements
Bachelor's and 7+ years of experience; relevant experience may be substituted for education
Demonstrated experience owning reliability (SLOs, observability, incident response) for production systems
Expert-level knowledge of at least one infrastructure-as-code tool (Terraform preferred)
Deep command of cloud infrastructure, containerization, and networking
Must be able to obtain and maintain a U.S. Public Trust / suitability determination
Tech Stack
AWS
Cloud
Kubernetes
Terraform
Benefits
Company-subsidized health, dental, and vision insurance