NVIDIA is a leading technology company focused on solving AI's toughest infrastructure problems. The Principal Architect role leads research and architectural direction for NVIDIA’s AI systems communication at scale, requiring expertise in high-performance networking and the ability to translate research into production-grade software.
Responsibilities:
- Setting the long-term technical vision for distributed AI communication systems—GPU-to-GPU, GPU-to-storage, and cross-node data movement
- Conducting original research and prototyping next-generation networking solutions over RDMA, NVLink, and GPUDirect
- Driving hardware-software co-optimization with GPU, DPU, NIC, and network switch. Investigating fundamental bottlenecks in communication runtimes for large-scale AI workloads (KV cache transfer, disaggregated prefill/decode, model parallelism)
- Integrating networking capabilities into AI serving stacks such as vLLM, SGLang, and TensorRT-LLM
- Publishing findings, representing NVIDIA in industry forums and standards bodies, and mentoring senior engineers across the organization