Nscale is a GPU cloud engineered for AI, providing high-performance infrastructure for AI start-ups and large enterprises. They are seeking a Staff HPC Systems Software Engineer to define the technical direction of their HPC platform, working across teams to enhance automation and operational approaches.
Responsibilities:
- Own and evolve the technical direction for a defined HPC systems domain, such as Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation
- Make architectural decisions that balance software quality, operational realities, customer needs, and long-term maintainability
- Define how proven Slurm implementations should be packaged, automated, and exposed as a service
- Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operating models across teams
- Act as the technical escalation point for the most complex issues within the domain
- Establish shared patterns and standards for automation, service lifecycle management, observability, reliability, and supportability across the HPC platform
- Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling
- Create reusable modules, automation, deployment patterns, and reference implementations that increase engineering leverage
- Identify and correct avoidable technical divergence, duplicated effort, and fragile operating models
- Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation, and production operations
- Lead technically critical initiatives spanning 2–4 teams or a defined HPC platform area
- Unblock delivery by clarifying technical direction and reducing ambiguity in complex system design problems
- Contribute hands-on where needed to de-risk or accelerate critical work
- Influence engineering teams without formal authority through strong judgement, design clarity, and practical solutions
- Partner with adjacent cloud-native software engineers so HPC implementations build on shared platform patterns rather than separate ones
Requirements:
- Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments
- Strong track record of writing maintainable, testable, and resilient software in Go, Python, or similar languages
- Proven ability to define technical direction across a domain spanning multiple teams or services
- Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concerns, and operational trade-offs
- Strong practical understanding of GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workload characteristics
- Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models
- Experience creating engineering leverage through standards, reusable patterns, shared tooling, and architectural clarity
- Strong judgement in balancing short-term delivery with long-term platform health and supportability
- Strong written and verbal communication skills, with the ability to align multiple teams around a coherent technical direction
- Experience with other schedulers or batch systems such as Kueue is valuable