Counterpart Health is transforming healthcare with its innovative primary care tool, Counterpart Assistant. They are seeking a Senior Site Reliability Engineer to support and improve their technology infrastructure, develop automation tools, and collaborate with various teams to maintain a scalable infrastructure platform.
Responsibilities:
- Build systems for declarative application and infrastructure lifecycle management, including continuous deployment, continuous integration, Kubernetes cluster management, and service/workload inventory
- Prioritize and troubleshoot infrastructure issues, minimizing downtime and responding to alerts efficiently
- Contribute to setting the direction of the Site Reliability Engineering (SRE) team, ensuring goals align with Counterpart Health’s company-wide objectives
- Foster a collaborative, high-performance culture that promotes motivation, innovation, and cross-disciplinary teamwork
- Streamline and automate infrastructure processes, including delivery pipelines and database changes
Requirements:
- 5+ years of programming experience and are proficient in at least one of the following languages: Python, Go, or Shell Scripting
- In-depth knowledge of containerization technologies and orchestration, such as Docker, Containerd, and Kubernetes, along with experience with CNCF-based technologies like Helm, gRPC, and Prometheus
- Experience with public cloud platforms such as GCP, Azure, or AWS
- Knowledgeable in networking fundamentals, including TCP/IP, UDP, firewalls, routing, DNS, and load balancing
- Experience with Linux system administration and a solid understanding of Linux design principles
- Understand key SRE concepts, such as monitoring, performance tuning, and automation
- Can work autonomously with limited guidance, proactively identifying and solving problems
- Excellent communication and collaboration skills, with the ability to work effectively with cross-functional teams and adapt to new challenges and evolving technologies