Cohere Health is a company focused on optimizing healthcare systems through clinical intelligence and AI-powered solutions. They are seeking a Site Reliability Engineer to ensure the availability and performance of their production healthcare systems, focusing on incident management and automation of operational tasks.
Responsibilities:
- Maintain the continuous uptime, scalability, and security of our AWS-hosted MERN applications and backend data architectures
- Manage, optimize, and troubleshoot event-driven architectures running on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts
- Monitor scheduled PySpark data workflows, execute standard operating procedures (SOPs) for large-scale data ingestion, and rapidly triage, rerun, or patch failed data processing jobs
- Participate in a collaborative on-call rotation to rapidly triage, debug, and mitigate live application outages and data flow bottlenecks
- Maintain strict HIPAA, SOC2, and HITRUST compliance profiles across all runtime environments, storage systems, and data pipelines handling Protected Health Information (PHI)
- Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark pipeline recovery steps
- Build specialized dashboards and alerts to monitor Node.js event loops, PySpark job execution stages, driver/worker memory leaks, and data pipeline throughput anomalies
- Lead blameless post-mortems for operational and data processing failures, translating system crashes into permanent structural fixes
Requirements:
- SaaS Platform Experience: Minimum of 3+ years of hands-on experience operating multi-tenant, cloud-hosted, or cloud-native SaaS platforms at scale
- AWS Cloud Engineering: Deep expertise operating AWS core services, specifically AWS Lambda, Amazon ECS/EKS, Amazon EMR or AWS Glue (for Spark), EC2, VPC networking, IAM permissions, and CloudWatch
- Automation & Data Languages: Professional competency in writing, debugging, and maintaining automation scripts and data tools using Python (including PySpark APIs) and Node.js
- Data Operations: Experience managing and troubleshooting distributed data orchestration pipelines, ETL tools, message queues (e.g., AWS SQS/SNS, RabbitMQ), or stream processing frameworks
- MERN Stack Operations: Deep understanding of the operational lifecycle of JavaScript/TypeScript applications, including memory management, asynchronous runtimes, and Node.js clustering
- Database Administration: Practical experience managing, sharding, indexing, and optimizing production-grade MySQL DB & Athena (RDS or self-hosted)
- Infrastructure as Code: Proven ability to deploy and maintain immutable infrastructure utilizing Terraform or OpenTofu
- Healthcare Experience: Minimum 1 year working within HIPAA-regulated environments. Direct experience securing data-at-rest and data-in-transit containing sensitive patient records is preferred
- Education & Experience: Minimum of 4 years of software/systems experience, with at least 1-2 years focused on live cloud operations and distributed data workflow management is preferred
- Crisis Management: Calm under pressure with a methodical approach to identifying and isolating PySpark driver OOM (Out of Memory) errors or data corruption during high-stress outages. Attention to detail and effective communications skills will be critical in working with clients and internal stakeholders is preferred