Tekmetric is a leading cloud-based platform for auto repair shops, dedicated to helping them thrive through innovative solutions. The Site Reliability Engineer will design and maintain scalable cloud infrastructure, ensuring system reliability and performance while collaborating with cross-functional teams.
Responsibilities:
- Design and implement scalable infrastructure: Architect and maintain reliable, scalable, and secure cloud infrastructure that supports positive user experiences and measurable business growth
- Monitor and optimize system performance: Develop and maintain monitoring, alerting, and incident response practices to ensure system reliability and performance at scale
- Automate everything: Create automated pipelines for deployment, testing, and infrastructure management to improve speed, consistency, and reliability across the organization
- Ensure high availability and disaster recovery: Implement and manage solutions for backup, disaster recovery, and failover processes to ensure business continuity
- Security and compliance: Apply best practices in security, monitoring, and compliance, ensuring that systems meet necessary requirements and regulations
- Collaboration: Work cross-functionally with development, data, product, and QA teams to improve application reliability and scalability
- Leadership and mentorship: Provide technical leadership, mentorship, and guidance to junior DevOps team members, fostering a culture of continuous learning and improvement
Requirements:
- 3+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related field, with deep knowledge of cloud environments (preferably AWS or GCP.)
- Hands-on experience with AWS (or similar cloud providers) and infrastructure as code (Terraform, etc.)
- Strong experience in automation tools
- Expertise in working with containerized environments like Docker and orchestration tools such as Kubernetes
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack)
- Proficiency in scripting languages like Python, Bash, or similar
- Experience with designing and optimizing Continuous Integration and Continuous Deployment (CI/CD) pipelines
- Strong communication skills and ability to work cross-functionally, solving complex technical challenges in a collaborative manner
- Ability to troubleshoot and resolve critical issues in high-pressure environments, maintaining composure and professionalism
- Experience with Infrastructure as Code tools like Terraform
- Familiarity with monitoring tools like Prometheus, Grafana, or the ELK stack
- Exposure to compliance and security best practices in cloud environments
- Experience coding in one or multiple programming languages such as Go, Java, Javascript