We are seeking a highly experienced Senior AI Infrastructure Engineer to deploy, operate, secure, and optimize NVIDIA DGX-based AI infrastructure supporting large-scale AI/ML training and inference workloads.
The ideal candidate will have strong hands-on expertise across NVIDIA DGX systems, Kubernetes, NVIDIA GPU Operator, InfiniBand, BlueField DPUs, NVLink/NVSwitch, and AI infrastructure automation. This is a deeply technical role requiring experience managing high-performance GPU clusters and secure, scalable Kubernetes environments.
AI/GPU: NVIDIA DGX, BasePOD, SuperPOD, NVLink, NVSwitch, NCCL, NVIDIA DCGM
Kubernetes: Kubernetes, GPU Operator, Helm, Kubeflow
Networking: InfiniBand, UFM, BlueField DPU, RoCE, RDMA
Automation: Python, Bash, YAML, Terraform, Ansible
DevOps: GitOps, CI/CD
Monitoring: Prometheus, Grafana
Storage: NFS, BeeGFS, Lustre