We are seeking a Principal AI Infrastructure Architect to design, build, and operate secure, scalable GPU-accelerated AI platforms. The ideal candidate will have deep expertise in Kubernetes, NVIDIA DGX infrastructure, InfiniBand networking, BlueField DPUs, and MLOps platforms.
This is a hands-on architecture role responsible for building production-grade AI infrastructure that supports large-scale ML training and inference workloads.
GPU Infrastructure: NVIDIA DGX, BasePOD, SuperPOD, NVIDIA AI Enterprise
Kubernetes: Kubernetes, GPU Operator, Helm, Custom Controllers, MIG
Networking: InfiniBand, UFM, BlueField DPU
MLOps: Kubeflow Pipelines, NVIDIA Triton Inference Server
Automation: Terraform, Ansible, Python, Bash, YAML
Monitoring: NVIDIA DCGM, Prometheus, Grafana