Dice is a company seeking a Lead Site Reliability Engineer specializing in Observability. The role involves designing and maintaining observability platforms and infrastructure, with a focus on tools like Splunk and Elasticsearch.
Responsibilities:
- Design, deploy, and operate enterprise observability platforms
- Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers
- Deploy and operate large-scale Elasticsearch clusters for log analytics and search
- Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry
- Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies
- Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions
- Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo
- Automate infrastructure using Terraform and configuration management tools
Requirements:
- 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps
- Hands-on experience administering Splunk Enterprise or Splunk Cloud
- Strong knowledge of Splunk SPL
- Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka
- Experience implementing metrics, logs, and traces as part of a modern observability strategy
- Experience with Terraform and Infrastructure as Code
- Programming experience in Python, Go, Ruby, or Bash
- Splunk certification
- Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies
- Experience supporting FedRAMP or regulated environments