Shadeform is a GPU cloud marketplace that aggregates compute across multiple cloud providers and data centers. They are seeking a Machine Learning Engineer to build an optimization and orchestration layer for managing ML workloads across a heterogeneous fleet of hardware.
Responsibilities:
- Build the optimization and orchestration logic that places and tunes ML workloads across a heterogeneous, multi-provider fleet
- Benchmark GPUs, interconnects, and driver stacks to understand real per-node performance, and feed that back into how the platform makes decisions
- Build the tooling, images, and reference stacks that take someone from a fresh instance to a running job without a day of setup
- Dig into performance and reliability problems that cross the line between the workload and the underlying hardware
- Turn what you learn into things that scale: defaults, playbooks, and platform improvements
- Take part in a weekly on-call rotation (24/7 coverage shared across the team)