NVIDIA is a leading technology company recognized for its groundbreaking developments in AI and cloud computing. They are seeking a Senior Software Engineer to join their DGX Cloud team, where the primary focus will be on building scalable automation solutions and integrating systems to enhance global cloud operations.
Responsibilities:
- Design and develop APIs to orchestrate and integrate operational workflows
- Build state management and workflow automation systems that streamline infrastructure lifecycle processes
- Collaborate across teams to codify business processes into scalable, self-measuring systems
- Develop extensible, schema-driven platforms for reducing manual toil and ensuring operational consistency
- Drive integrations with container orchestration tools like Kubernetes and observability systems such as Prometheus, OpenTelemetry, Grafana
- Optimize the reliability and efficiency of cloud operations through automated workflows and telemetry systems
- Lead and ship impactful technical projects, ensuring quality and scalability at every stage
Requirements:
- 8+ years of industry experience with a Bachelor's or Master's degree (or equivalent experience), or 2+ years with a PhD
- Expertise in designing, building, and operating services in a high reliability environment
- Proficiency in programming languages such as Go, Java, or Python
- Strong understanding of cloud infrastructure (AWS, GCP, Azure) and container technologies like Docker and Kubernetes
- Experience with high-scale distributed systems, including architectural patterns for APIs and data pipelines
- Outstanding communication and collaboration skills, with a focus on solving complex operational challenges
- A passion for automating manual processes and driving system efficiency
- A track record of designing workflow orchestration systems for large-scale infrastructure
- Proven experience in reducing operational inefficiencies through automation and integration
- Strong debugging and problem-solving skills in distributed environments
- Prior experience or strong familiarity with the operational aspects of the NVIDIA AI/ML software stack (e.g., CUDA, cuDNN, containerization)