Zscaler accelerates digital transformation to enhance agility, efficiency, resilience, and security for its customers. They are seeking a Principal Software Development Engineer to build backend systems, create microservices, and ensure efficient data ingestion and querying while collaborating with multiple teams to deliver high-quality results.
Responsibilities:
- Architect and deliver scalable microservices on EKS (GoLang): design, build, and operate production services that meet uptime, latency, and scale goals
- Build event-driven pipelines with Kafka and MQTT: implement asynchronous workflows that are resilient, observable, and easy to evolve as product needs change
- Design high-performance data models and access patterns: optimize ingest and querying across RDS MySQL, Neo4J, and ElasticSearch for correctness, speed, and cost efficiency
- Harden production readiness and operational excellence: establish monitoring, alerting, tracing, and on-call standards; drive root-cause fixes and reliability improvements
- Lead cross-team delivery and technical direction: align stakeholders, break down complex work, and execute quickly while maintaining a high bar for quality and security
Requirements:
- Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain
- 12+ years of experience in backend engineering designing and shipping highly scalable microservices in production developed in Go
- Proven experience with event-driven, asynchronous architectures using messaging/streaming platforms such as Kafka
- Strong data engineering and persistence skills building efficient ingest and query patterns across relational databases (MySQL) plus at least one of ElasticSearch or Neo4J
- Hands-on experience running services on Kubernetes (ideally EKS), including deployment, troubleshooting, and operational ownership
- Experience leveraging AI-driven development tools, predictive performance modeling, or intelligent microservice telemetry to optimize high-throughput distributed systems
- Proven ability to profile and optimize distributed systems end-to-end (latency, throughput, resource usage), including performance tuning across services, databases, and messaging
- Strong track record designing debuggability/observability for distributed applications alongside experience defining, operationalizing, and driving product decisions based on SLIs/SLOs