Cerebras builds AI hardware and inference systems designed to deliver high-performance model training and serving. The Software Engineer, GPU Inference will productionize and optimize the GPU serving stack across custom inference APIs, vLLM, PyTorch, ROCm, and AMD GPU infrastructure, improving reliability, correctness, observability, and performance.