NVIDIA develops accelerated computing and AI infrastructure, including its Base Command Manager platform for deploying and operating AI data centers. The Senior Site Reliability Engineer will support large-scale GPU platform deployments and operations, handle cluster incidents, implement product features, and validate Slurm and Kubernetes configurations for performance, scalability, and resilience.