Coterie Insurance is on a mission to build a world-class team to enhance commercial insurance through innovative technology. They are seeking a Senior Site Reliability Engineer to manage cloud infrastructure, improve CI/CD pipelines, and collaborate with development teams to ensure reliable software delivery.
Responsibilities:
- Manage and maintain cloud infrastructure on Azure, including Azure Kubernetes Service (AKS) clusters and supporting resources
- Build, improve, and maintain CI/CD pipelines using GitHub Actions to support reliable and repeatable deployments
- Own and enhance our Grafana implementation; designing dashboards, configuring alerts, and supporting incident management workflows
- Monitor system health, triage incidents, and drive root cause analysis to prevent recurrence
- Collaborate with development teams to define and track SLIs, SLOs, and error budgets that align with business goals
- Contribute to infrastructure-as-code practices using Pulumi
- Identify and resolve reliability risks through capacity planning, performance tuning, and proactive system improvements
- Participate in an on-call rotation to support production systems and respond to incidents
- Document runbooks, operational procedures, and architectural decisions to support team knowledge sharing
Requirements:
- 5+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure role
- 3+ years experience working with infrastructure as code
- 2+ years of experience architecting CI/CD pipelines and cloud-based infrastructure
- Strong hands-on experience with Azure Cloud services and resource management
- Kubernetes and AKS administration, including deployments, networking, and troubleshooting
- GitHub Actions for CI/CD pipeline development and maintenance
- 3+ experience with Grafana or similar tooling, including dashboard creation, alerting configuration, and incident management
- Hands-on experience with Prometheus, Loki, or other observability tools in the Grafana ecosystem
- Proficiency in at least one scripting or programming language such as Python or Bash
- Understanding of networking fundamentals, DNS, load balancing, and container orchestration concepts
- Strong analytical and communication skills; able to diagnose complex system issues and clearly communicate findings
- Demonstrated ability to collaborate across teams and contribute to a culture of reliability
- Experience working in an agile environment with modern DevOps practices
- Experience working at a startup or in a fast-paced, cross-functional environment
- Familiarity with the insurance industry or other regulated sectors
- Familiarity with service mesh technologies (e.g., Istio)