Convective is a company focused on enhancing weather forecasts through innovative technology, including a global network of smart weather balloons. They are seeking a Machine Learning Infrastructure Engineer to streamline operations, manage data pipelines, and ensure the reliability of their AI weather models.
Responsibilities:
- Research to Operations pipelines — Our models serve real-time forecasts to customers with strict latency requirements. You'd own uptime end-to-end: build health monitoring, improve logging, diagnose failures across nodes
- Inference scaling & compute strategy — We have an on-prem cluster but also use cloud providers, especially for production deployments. You'd evaluate cost/performance tradeoffs across cloud options as we scale, and also help manage growing on-prem resources for compute and storage
- Data pipelines & upstream reliability — Weather data comes from dozens of sources (satellites, government agencies, our own balloon observations) with varying schedules, incomplete documentation and sometimes failing or changing quality. You'd build pipelines for training and realtime data that gracefully handle upstream delays, do QC checks on data, and add logging and alerting for a zoo of edge cases
- Training infrastructure — Make distributed training runs reliable. They die from silent OOMs, network faults, and storage issues. Build monitoring, auto-recovery, and job scheduling so researchers can launch experiments with less need for babysitting them