Nuance Labs is a research company focused on building advanced AI avatars with emotional intelligence. They are seeking a Member of Technical Staff to design and operate large-scale data pipelines for multimodal training data, ensuring high throughput and data quality.
Responsibilities:
- Design, build, and operate large-scale data pipelines for ingestion, processing, filtering, and curation of multimodal training data (video, audio, text)
- Take research-grade data processing code and turn it into robust, production-level pipelines — quickly and without losing correctness
- Optimize pipeline throughput and efficiency at scale; identify and eliminate bottlenecks across compute, I/O, and storage
- Build and maintain data quality systems — deduplication, filtering, validation, and quality scoring at scale
- Manage petabyte-scale datasets: storage architecture, versioning, lineage tracking, and cost efficiency
- Work closely with researchers to understand data requirements and translate them into scalable processing systems
- Build tooling and infrastructure that makes the research team faster — efficient data access, reproducible processing, and fast iteration loops