Develop, extend, and optimize inference runtime configurations and integrations across vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, and Triton
Write Python-based tooling and automation for model onboarding, serving configuration, performance benchmarking, and deployment pipelines
Build and maintain Kubernetes-native model serving infrastructure using KServe, KNative, and OpenShift AI — including custom serving runtimes and inference graphs
Implement and tune inference performance optimizations — continuous batching, speculative decoding, prefix caching, concurrency control, autoscaling policies, and disaggregated prefill/decode pipelines
Develop Helm charts, operators, and Kustomize overlays for deploying and managing inference workloads on OpenShift/OCP
Integrate inference platforms with GPU workload orchestrators (Run:AI or similar) — automating project provisioning, quota management, and workload scheduling
Build observability and testing harnesses — load testing frameworks, latency/throughput profiling scripts, and regression test suites for inference stack upgrades
Partner with AI/ML teams to productionize new models, defining serving architectures, resource requirements, and SLA targets

7+ years in software engineering or platform engineering (work experience, training, military experience, or education)
5+ years of programming experience in Python with experience building production systems
Experience with Inference frameworks, such as vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or Triton Inference Server
Experience with Kubernetes-native ML serving, such as KServe, KNative, Seldon, or OpenShift AI
Experience with Inference optimization, (Continuous batching, speculative decoding, KV-cache management, prefix caching, quantization-aware serving (FP8, AWQ, GPTQ), or tensor parallelism configuration)
Experience with Container platform development, (Writing Helm charts, operators, or custom controllers for OpenShift, GKE, or EKS)
Experience with GPU workload orchestration, (Run:AI, Kueue, Volcano — scripting workload automation, quota management, or scheduler integrations)
Experience with Performance and load testing, (Building benchmarking tools for token throughput, time-to-first-token, batch latency, and autoscaling behavior)
Familiarity with NVIDIA GPU fundamentals (CUDA, MIG, NCCL), experience contributing to open-source inference projects, or background in ML observability tooling (Prometheus, Grafana, Arize)

Principal Engineer – Gen AI Platform Inferencing Engineering

Key skills