Dice is seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-scale data infrastructure environments across AWS and on-premises Hadoop platforms. The role focuses on ensuring the availability, scalability, security, and operational excellence of mission-critical data platforms while driving continuous improvements through modern SRE practices.
Responsibilities:
- Maintain and support highly available, scalable, and secure data infrastructure platforms across AWS and on-premises environments
- Drive operational excellence through automation of repetitive tasks, incident reduction, and proactive reliability improvements
- Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency
- Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives
- Participate in on-call rotations and partner with teams across the US and India to provide 24x7 operational support. US support is aligned to Pacific Time zone
- Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise
Requirements:
- Strong SRE mindset with a proven focus on reliability, availability, performance optimization, incident management, and operational excellence
- Experience delivering services within defined SLAs and ensuring timely resolution of production issues
- Expertise in troubleshooting complex distributed systems and identifying root causes quickly and effectively
- Deep hands-on experience with AWS services, including EMR, EKS, MSK, Athena, Glue, IAM, Amazon S3, VPC, AWS networking and security services
- Strong understanding of cloud-native architectures, scalability, and infrastructure resilience