Dice is seeking a Staff Machine Learning Engineer to own the execution layer of intelligence. This remote role involves translating research into production-grade Machine Learning systems, ensuring models are trainable, deployable, and performant under real-world constraints.
Responsibilities:
- Own end-to-end Machine Learning (ML) system execution: data pipelines, training workflows, evaluation systems, inference architecture, and deployment
- Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation
- Architect and operate scalable inference systems, balancing latency, cost, and reliability
- Design and maintain data systems for high-quality synthetic and real-world training data
- Implement evaluation pipelines covering performance, robustness, safety, and bias, in partnership with research leadership
- Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies
- Collaborate closely with application engineering to integrate Machine Learning (ML) systems cleanly into backend, mobile, and desktop products
- Make pragmatic trade-offs and ship improvements quickly, learning from real usage
- Work under real production constraints: latency, cost, reliability, and safety
Requirements:
- Experience building or shipping real Machine Learning (ML) systems used by people, not just demos
- Artificial Intelligence (AI) experience required
- Experience working with large models and understanding their failure modes
- Experience writing strong, production-grade code
- You are self-directed, pragmatic, and take full ownership of outcomes
- Experience communicating clearly and collaborate well in small, high-trust teams
- Tech Stack: GPU-based training and inference system, JAX, Python, and PyTorch