RadixArk is an infrastructure-first company focused on democratizing AI technology. They are seeking a Member of Technical Staff, Developer Technology to enhance GPU performance for AI workloads and contribute to their open-source systems.
Responsibilities:
- Accelerate AI workloads. Profile and optimize GPU performance for real production workloads on current and next-generation hardware, root-causing bottlenecks from kernels to distributed multi-node systems
- Go deep in one or two focus areas. The team collectively covers the full stack; each engineer specializes in one or two tracks:
- Inference performance: engine tuning, benchmarking, long-context and multi-turn optimization, parallelism strategy, production debugging
- Kernels and model/hardware enablement: custom CUDA/ROCm/Triton kernels, low-precision quantization, day-0 support for new models on new silicon
- Speculative decoding: draft-model training, acceptance-rate tuning, cross-platform kernel adaptation
- Training systems: RL post-training with Miles, FP8 training, elasticity, long-rollout and long-context efficiency
- Partner directly with the ecosystem. Turn ambiguous, high-stakes problems from expert engineers at our key partners into concrete wins, clear technical guidance, and reproducible cookbooks
- Enhance SGLang and Miles. Feed user-driven improvements back into our open-source systems and roadmap, so every win compounds across the ecosystem