Software Technology Inc. is seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-scale data infrastructure environments. The role focuses on ensuring the availability, scalability, security, and operational excellence of mission-critical data platforms while driving continuous improvements through Infrastructure as Code and modern SRE practices.
Responsibilities:
- Maintain and support highly available, scalable, and secure data infrastructure platforms across AWS and on-premises environments
- Drive operational excellence through automation of repetitive tasks, incident reduction, and proactive reliability improvements
- Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency
- Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives
- Participate in on-call rotations and partner with teams across the US and India to provide 24x7 operational support. US support is aligned to Pacific Time zone
- Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise
Requirements:
- Site Reliability Engineering (SRE)
- Strong SRE mindset with a proven focus on reliability, availability, performance optimization, incident management, and operational excellence
- Experience delivering services within defined SLAs and ensuring timely resolution of production issues
- Expertise in troubleshooting complex distributed systems and identifying root causes quickly and effectively
- Deep hands-on experience with AWS services, including EMR, EKS, MSK, Athena, Glue, IAM, Amazon S3, VPC, AWS networking and security services
- Strong understanding of cloud-native architectures, scalability, and infrastructure resilience
- Extensive operational experience managing Hadoop clusters, with a strong focus on administration, platform maintenance, and automation of day-to-day operational activities
- Experience supporting both AWS-based data platforms and on-premises Cloudera CDH/CDP environments
- Solid understanding of Kerberos authentication and security implementation within Hadoop ecosystems
- Hands-on experience with Apache Spark
- Hands-on experience with Apache Iceberg
- Big Data platform architecture
- Performance tuning and optimization
- Strong Linux administration and operational support experience
- Expertise in user and access management, system configuration, customization, and platform administration
- Hands-on experience with monitoring and observability platforms such as AWS CloudWatch, Datadog, PagerDuty, Similar enterprise monitoring solutions
- Proven success improving alert quality, reducing false positives, and minimizing alert fatigue
- Excellent debugging and troubleshooting skills across infrastructure, applications, Spark workloads, and Iceberg environments
- Strong understanding of Java application administration, including JVM tuning, Thread dump analysis, Heap dump analysis, JVM parameters, Application log analysis and troubleshooting
- Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases
- Strong experience with Terraform
- Infrastructure as Code (IaC)
- CI/CD pipeline implementation and automation
- Platform engineering best practices
- Current hands-on experience applying AI technologies to Data Infrastructure and SRE operations
- Demonstrated ability to design and implement Agentic AI solutions
- AI-assisted operational workflows
- Intelligent automation for routine SRE activities
- Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity