Unity is the world’s leading game engine, powering play for more than 3 billion consumers each month. They are seeking a Senior ML engineer to design and evolve Unity Vector’s online model inference platform, focusing on building reliable infrastructure for serving machine learning models in production and optimizing inference performance.
Responsibilities:
- Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or similar distributed serving frameworks
- Develop infrastructure that supports distributed training workflows using technologies such as Pytorch, Ray Data, and Ray Train, etc
- Integrate ML pipelines with workflow orchestration systems (e.g., Flyte, Airflow, or similar) to enable reliable multi-stage training workflows
- Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning
- Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring
- Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency
- Improve the reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation
- Lead architectural improvements that make the online ML platform more robust, user-friendly, scalable, and cost-efficient
Requirements:
- Experience building and operating production-grade online ML inference systems, such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems
- Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems
- Experience optimizing inference workloads using techniques such as dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning
- Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability
- Strong programming skills in Python, with practical experience working on production ML systems and high-scale services
- Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management
- Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback
- Strong systems thinking, with the ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs in online systems
- Proven ability to lead technical direction and influence architectural decisions across teams without formal authority