NVIDIA is a leader in AI technology, seeking a Software Engineer specializing in Deep Learning Inference. The role involves designing, building, and optimizing GPU-accelerated software for advanced AI applications, as well as contributing to high-performance open-source frameworks.
Responsibilities:
- Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI
- Scale performance of DL models across different architectures and types of NVIDIA accelerators
- Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions
- Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions