Research, prototype, and develop state-of-the-art computer vision and deep learning models spanning detection, segmentation, tracking, and vision-language tasks.
Translate recent research into working implementations, reproducing baselines, running ablations, and quantifying practical tradeoffs.
Own the full data lifecycle for visual tasks: dataset curation, annotation strategy, augmentation, synthetic data generation, and quality analysis.
Collaborate closely with engineering and product teams to move models from prototype to production, including optimization for edge or latency-constrained environments.
Communicate methods, results, and tradeoffs clearly to both technical and non-technical stakeholders.
Requirements
Master’s or PhD in Computer Science, Electrical Engineering, Robotics, Machine Learning, or a related field, or a Bachelor’s with equivalent research or industry experience.
Strong foundation in modern computer vision and deep learning: image classification, object detection, segmentation, and tracking across CNNs, Vision Transformers, vision-language models, and foundation models.
Solid grasp of deep learning fundamentals: supervised and self-supervised learning, representation learning, optimization, loss design, and rigorous model training and evaluation.
Demonstrated ability to read, implement, and build on recent research papers, turning literature into working prototypes.
Hands-on experience with real-world visual data: dataset curation, annotation, augmentation, synthetic data, data-quality analysis, and low-data or noisy-data settings.
Strong applied mathematics (linear algebra, probability, statistics, optimization) and proficiency in Python with PyTorch or equivalent deep learning framework.
Strong collaboration and communication skills — able to work across research, engineering, and product, and present technical work clearly to mixed audiences.