Ad Hoc LLC is a technology company that empowers organizations to deliver scalable, impactful digital services. As a Site Reliability Engineer, you will help ensure the availability, performance, and reliability of a large federal enterprise cloud platform while improving the platform's reliability practices.
Responsibilities:
- Monitoring platform health and supporting service level objectives (SLOs), service level indicators, and error budgets
- Building and maintaining observability tooling, including metrics, logging, alerting, and dashboards
- Participating in on-call rotations and incident response, helping restore service and reduce time to recovery
- Contributing to blameless postmortems and driving follow-up actions
- Automating repetitive operational tasks to reduce toil
- Supporting capacity planning and performance tuning across cloud infrastructure (AWS) and Kubernetes (Amazon EKS)
- Implementing reliability improvements as infrastructure as code (Terraform)
- Working with government partners and application teams to meet security, SLA, and performance requirements
- Supporting recruiting efforts by evaluating exercises and assisting with interviews
Requirements:
- Bachelor's and 5+ years of experience; relevant experience may be substituted for education
- Experience with monitoring and observability tooling and on-call operations
- Proficient with at least one infrastructure-as-code tool (Terraform preferred)
- Background in key DevOps concepts: containerization, networking, and cloud infrastructure
- Must be able to obtain and maintain a U.S. Public Trust / suitability determination
- Prior experience with the Department of Veterans Affairs
- Experience with Kubernetes (Amazon EKS) and AWS in production
- Familiarity with SLO-based reliability practices and error budgets
- Relevant certifications (e.g., AWS, Certified Kubernetes Administrator)