Netflix is the world's leading streaming entertainment service, and they are seeking a Senior Engineer with expertise in distributed model training for their AI Platform. The role involves designing and building a platform for large-scale machine learning model training and optimizing systems for performance and reliability.
Responsibilities:
- Design and build the platform that powers large-scale machine learning model training, fine-tuning, model transformation and evaluations workflows and use cases from the entire company
- Co-design and optimize the systems and models to scale up and increase the cost-effectiveness of machine learning model training
- Design easy-to-use APIs and interfaces for experienced ML practitioners, as well as non-experts to easy access the training platform
- Design, build, and operate platform infrastructure, libraries, and SDKs for large-scale model training. Enable reliable and efficient training workflows for foundation models and generative AI models of all sizes
- Diagnose and optimize the performance of large distributed training jobs, including GPU utilization, memory efficiency, communication overhead, data loading, checkpointing, fault tolerance, and cluster utilization
- Experience with cloud computing providers, preferably AWS
- Comfortable with ambiguity and working across multiple layers of the tech stack to execute on both 0-to-1 and 1-to-100 projects
- Adopt and promote best practices in operations, including observability, logging, reporting, and on-call processes to ensure engineering excellence
- Excellent written and verbal communication skills
- Comfortable working in a team with peers and partners distributed across (US) geographies & time zones
- Understand modern and real-world Machine Learning model development workflows and experience partnering closely with ML modeling engineers. Lead technical design reviews, facilitate cross-functional discussions, communicate tradeoffs clearly, and align stakeholders on platform direction and execution priorities
- Familiarity with cloud-based AI/ML services (e.g., SageMaker, Bedrock, Databricks, OpenAI, etc.)
- Familiarity with distributed training performance analysis tools and techniques, such as PyTorch Profiler, NVIDIA Nsight Systems, GPU telemetry, communication profiling, or cluster-level utilization analysis
- Experience with large-scale distributed training and different parallelism techniques for scaling up training, such as FSDP and tensor/pipeline parallelism
- Expertise in the area of Generative AI, specifically when it comes to training foundation models, fine-tuning them, and distilling them to smaller models
Requirements:
- Design, build, and operate platform infrastructure, libraries, and SDKs for large-scale model training. Enable reliable and efficient training workflows for foundation models and generative AI models of all sizes
- Diagnose and optimize the performance of large distributed training jobs, including GPU utilization, memory efficiency, communication overhead, data loading, checkpointing, fault tolerance, and cluster utilization
- Experience with cloud computing providers, preferably AWS
- Comfortable with ambiguity and working across multiple layers of the tech stack to execute on both 0-to-1 and 1-to-100 projects
- Adopt and promote best practices in operations, including observability, logging, reporting, and on-call processes to ensure engineering excellence
- Excellent written and verbal communication skills
- Comfortable working in a team with peers and partners distributed across (US) geographies & time zones
- Understand modern and real-world Machine Learning model development workflows and experience partnering closely with ML modeling engineers. Lead technical design reviews, facilitate cross-functional discussions, communicate tradeoffs clearly, and align stakeholders on platform direction and execution priorities
- Familiarity with cloud-based AI/ML services (e.g., SageMaker, Bedrock, Databricks, OpenAI, etc.)
- Familiarity with distributed training performance analysis tools and techniques, such as PyTorch Profiler, NVIDIA Nsight Systems, GPU telemetry, communication profiling, or cluster-level utilization analysis
- Experience with large-scale distributed training and different parallelism techniques for scaling up training, such as FSDP and tensor/pipeline parallelism
- Expertise in the area of Generative AI, specifically when it comes to training foundation models, fine-tuning them, and distilling them to smaller models