Cracker Barrel is a well-established company known for its welcoming atmosphere and commitment to service. They are seeking a Site Reliability Engineer to design and maintain reliability practices across digital platforms, provide production support, and collaborate with various teams to ensure operational excellence.
Responsibilities:
- Design, implement, and maintain reliability practices that improve availability, performance, scalability, resiliency, and operational maturity across digital platforms and supporting services
- Monitor production systems using observability tools, dashboards, logs, metrics, traces, alerts, and synthetic monitoring to identify issues before they impact customers or associates
- Provide Tier-2 and Tier-3 production support for customer-facing and internal digital applications, including web, mobile, commerce, CMS, APIs, integrations, and cloud-hosted services
- Lead and participate in incident response, root-cause analysis, problem management, post-incident reviews, and follow-up actions that reduce recurrence and improve service reliability
- Develop automation, scripts, runbooks, self-healing processes, and operational tools that reduce manual effort, accelerate recovery, and improve consistency across environments
- Partner with development teams to improve CI/CD pipelines, deployment readiness, release validation, rollback procedures, feature monitoring, and environment stability
- Collaborate with infrastructure, cloud, security, architecture, QA, and vendor teams to ensure systems meet company standards for security, privacy, compliance, resiliency, and operational support
- Define and track service health indicators such as availability, latency, error rates, capacity, incident trends, deployment quality, and other reliability metrics
- Create and maintain technical documentation, operational support guides, escalation paths, production readiness checklists, and disaster recovery procedures
- Understand and comply with all company privacy, security, accessibility, change management, and technology standards
Requirements:
- 3–5+ years of experience in site reliability engineering, DevOps, cloud operations, production support, systems engineering, software engineering, or a related technology operations role
- Experience supporting high-availability web, mobile, commerce, API, integration, or cloud-hosted application environments
- Hands-on experience with monitoring, logging, alerting, incident management, root-cause analysis, CI/CD pipelines, Git-based workflows, and release support
- Bachelor's degree in Computer Science, Computer Information Systems, Software Engineering, Information Technology, or a related discipline is preferred; equivalent experience or training may be considered
- Experience with cloud platforms, containers, infrastructure automation, scripting, APIs, microservices, content management systems, or restaurant/retail technology environments preferred