Abacus Insights is transforming how data works for health plans, making healthcare data usable for better decision-making. The Principal Data Engineer will design and scale the enterprise data platform, architecting complex data integration solutions and ensuring data engineering excellence across the organization.
Responsibilities:
- Architect Enterprise‑Scale Data Solutions: Design, build, and evolve high‑volume batch and real‑time data pipelines using PySpark, SparkSQL, Databricks Workflows, and distributed processing frameworks
- Own Platform‑Level Integrations: Develop end‑to‑end ingestion and transformation frameworks integrating Databricks, Snowflake, AWS services (such as S3, SQS, Lambda), and external data provider APIs, with a strong focus on data quality, lineage, and schema evolution
- Lead Technical Design for Clients: Serve as the technical lead for complex client implementations, defining highly available, fault‑tolerant architectures across multi‑account cloud environments
- Translate Business Needs into Architecture: Convert complex business and regulatory requirements into scalable technical designs, detailed specifications, and reusable engineering patterns
- Set Engineering Standards: Establish and champion best practices across CI/CD, code quality, testing, orchestration, monitoring, logging, and observability for data platforms
- Ensure Security & Compliance: Design and implement security‑first data solutions, including RBAC, encryption, PHI handling, auditability, and alignment with HIPAA and SOC 2 requirements
- Optimize Performance & Cost: Profile and tune compute workloads, cluster configurations, partitioning strategies, indexing, and caching across Databricks and Snowflake environments
- Provide Technical Mentorship: Mentor senior and junior engineers, conduct design and code reviews, and raise the overall technical bar across teams
- Produce Technical Artifacts: Create clear documentation, including architecture diagrams, runbooks, and operational standards that support scalable delivery
Requirements:
- 7+ years of hands‑on experience designing and operating large‑scale, distributed data systems in cloud‑based environments
- Expert‑level proficiency in Python, SQL, and PySpark, including performance‑optimized distributed transformations
- Proven experience building and operating production‑grade ETL/ELT pipelines using Databricks, Airflow, or similar orchestration frameworks
- Strong working knowledge of AWS‑based data services (e.g., S3, SQS, Lambda, IAM) or equivalent cloud technologies
- Experience working with dbt, Delta Lake, Kafka, or event‑driven architectures in modern data platforms
- Hands‑on experience with Snowflake or other cloud data warehouses, including schema design and performance optimization
- Demonstrated ability to design scalable, resilient systems requiring specialized knowledge of distributed computing and cloud‑scale data processing
- Working experience with healthcare data domains such as claims, eligibility, provider, or clinical datasets
- Ability to clearly communicate complex technical concepts to both technical and non‑technical partners
- Bachelors or Masters degree in Computer Science, Engineering, Data Science, or a related field
- Experience in large‑scale healthcare payer or analytics environments
- Exposure to machine learning or advanced analytics workflows
- Familiarity with DevOps or platform engineering practices applied to data systems
- Experience developing reusable data frameworks or shared platform capabilities