NextAmp LLC is a digital modernization services company focused on transforming the insurance industry. They are seeking an experienced DevOps / Site Reliability Engineer (SRE) to build scalable infrastructure, improve CI/CD pipelines, and ensure the reliability and performance of mission-critical applications. The ideal candidate will have strong expertise in cloud infrastructure, automation, and observability.
Responsibilities:
- Design, implement, and maintain scalable and secure cloud infrastructure on AWS
- Build, automate, and optimize CI/CD pipelines using GitHub Actions and other modern DevOps tools
- Lead migration of Apache Airflow v2 to Airflow v3 on the Astronomer platform
- Design, develop, optimize, and troubleshoot Airflow DAGs for production data workflows
- Deploy, manage, and maintain Kubernetes workloads on Amazon EKS
- Implement Infrastructure as Code (IaC) using Terraform , CloudFormation, or similar tools
- Implement and maintain observability solutions using the Grafana Stack , including Grafana Alloy and Grafana Beyla
- Utilize eBPF technologies for deep infrastructure and application observability
- Build and maintain automated testing infrastructure supporting CI/CD pipelines
- Monitor application and infrastructure health, respond to production incidents, and perform Root Cause Analysis (RCA)
- Automate operational processes to improve deployment reliability and reduce manual effort
- Collaborate closely with software engineering teams to improve deployment processes, platform stability, and system reliability
- Implement logging, monitoring, alerting, and performance optimization best practices
- Ensure high availability, scalability, security, and operational excellence across production environments
Requirements:
- 5–8 years of experience in DevOps or Site Reliability Engineering
- Strong programming experience in Python
- Hands-on experience migrating Apache Airflow v2 to Airflow v3
- Experience working with the Astronomer platform
- Strong knowledge of Airflow DAG development, scheduling, monitoring, debugging, and optimization
- Experience building and managing CI/CD pipelines using GitHub Actions
- Experience designing and maintaining CI testing infrastructure
- Strong experience with Docker and Kubernetes
- Hands-on experience managing workloads on Amazon EKS (Elastic Kubernetes Service)
- Strong knowledge of Infrastructure as Code using Terraform, CloudFormation, or similar tools
- Experience with AWS cloud services and cloud-native architectures
- Hands-on experience with the Grafana Stack, including Grafana, Grafana Alloy, and Grafana Beyla
- Deep understanding of eBPF for Linux observability, networking, and performance monitoring
- Experience implementing monitoring, logging, tracing, and alerting solutions
- Strong Linux administration, networking, and security fundamentals
- Strong troubleshooting skills with production incident management experience
- Experience with Helm, ArgoCD, or FluxCD
- Experience with OpenTelemetry and distributed tracing
- Knowledge of Prometheus, Loki, Tempo, and modern observability platforms
- Experience with service mesh technologies such as Istio or Linkerd
- Familiarity with secrets management solutions such as HashiCorp Vault or AWS Secrets Manager
- Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and reliability engineering
- Experience supporting highly available, distributed production systems
- Experience implementing cloud governance and infrastructure cost optimization
- Knowledge of disaster recovery, backup, and business continuity strategies