PointClickCare is a leading health tech company focused on empowering providers to deliver exceptional care. They are seeking a Senior Applied Research Data Engineer who will work at the intersection of data engineering and applied research to transform complex healthcare data into AI-ready research assets.
Responsibilities:
- Build and own reusable gold-layer data products that power AI, machine learning, and generative AI research
- Transform structured, semi-structured, and unstructured healthcare data into trusted, model-ready datasets
- Investigate and document complex business logic by analyzing source systems, stored procedures, application code, and stakeholder workflows
- Partner directly with researchers to design datasets for experimentation, evaluation, and model training
- Create semantic data definitions, lineage documentation, provenance records, and data quality frameworks that enable reproducible research
- Develop point-in-time-correct datasets, feature sets, and evaluation corpora for classical ML and generative AI workloads
- Support advanced AI data preparation techniques including programmatic labeling, weak supervision, synthetic data generation, and research dataset curation
- Serve as a bridge between domain experts, researchers, and engineering teams, turning tacit knowledge into durable data assets
Requirements:
- 5+ years building production data systems, with at least 2 supporting ML or AI workloads
- Track record of learning complex new data domains quickly, through reading source code, interviewing experts, and building durable artifacts others rely on
- Advanced Python, SQL, and PySpark/Databricks for working with large, messy data. Expert SQL specifically: comfortable reading complex stored procedures and reverse-engineering business logic from queries
- Databricks ecosystem depth: Delta Lake, Unity Catalog, Spark/PySpark tuning, MLflow
- AI domain literacy: working understanding of embeddings, tokenization, feature engineering, point-in-time correctness, train/validation/test splits, data drift, and the differences between what classical ML and generative models need from data
- Data wrangling across modalities: transforming unstructured content (text, PDFs, transcripts, logs) and structured tabular data into clean, model-ready forms
- AI-friendly data formats (Parquet, Hugging Face datasets) and storage layout decisions — partitioning, sharding, caching, that keep researcher workflows responsive in Azure, AWS or other working environments
- Data quality, filtering, and synthesis pipelines: support for programmatic labeling and weak supervision (e.g. Snorkel or equivalent), near-duplicate detection (MinHash/LSH), content and quality filters, LLM-API-driven synthetic data generation
- Pipeline orchestration (e.g. a la Airflow, Databricks Workflows, Dagster, or Prefect) and dataset versioning including Unity Catalog and feature-store support
- Experience handling regulated or sensitive data under controlled access (HIPAA or equivalent). Familiarity with general de-identification concepts
- Git-based version control and CI/CD for data and code
- Strong written documentation. Skill in eliciting requirements and tacit knowledge from technical and non-technical experts
- Bachelor's degree in computer science, data science, engineering, statistics, or related field. Equivalent practical experience considered
- Hands-on EHR data experience, ideally in skilled nursing, long-term care, post-acute care, or senior living
- Working knowledge of clinical terminologies (ICD-10, SNOMED CT, LOINC) and data standards (HL7v2, FHIR, CCDA)
- dbt for transformation and testing
- Familiarity with training-side ML frameworks (e.g. PyTorch) sufficient to debug data-side bottlenecks; experience supporting LLM or foundation-model training or fine-tuning data pipelines
- Clinical NLP, OCR, document parsing, or ASR / transcript pipeline experience
- Data lineage and catalog tools
- Prior experience embedded inside an AI or ML research team
- Master's degree in a relevant quantitative or computer science field