Job Title: Databricks Data Lake Engineer
Location: New York, NY - Jersey City, NJ
Role Overview:
The Databricks Data Lake Engineer will design, build, and optimize large scale data pipelines across the full lifecycle of data ingestion, migration, curation, and consumption within a modern lakehouse architecture. This role requires hands on expertise with Databricks, Delta Lake, Spark, and cloud-native data platforms, along with the ability to collaborate effectively with business stakeholders, client, architects, and AI/ML teams.
Key Responsibilities:
- Data Ingestion - Build scalable ingestion pipelines using Spark, Autoloader, LakeFlow, Informatica, Delta Live Tables, and cloud-native connectors (Kafka, REST, ODBC, CDC - Change Data Capture).
- Data Migration - Lead migration of legacy data warehouses, Hadoop clusters, or on prem systems into S3/Delta Lake.
- Data Curation - Implement bronze silver gold architecture, enforce quality checks, schema evolution, and governance.
- Data Consumption - Deliver curated datasets for BI, analytics, dashboards, and downstream applications.
- AI/ML Enablement - Partner with data scientists to prepare feature stores, optimize ML-ready datasets, and support model deployment workflows.
- Develop and maintain CI/CD pipelines for Databricks jobs, notebooks, and workflows.
- Optimize Spark/SQL/Python jobs for performance, cost efficiency, and reliability.
- Implement security, governance, and compliance using Unity Catalog, data lineage, and access controls.
- Collaborate with cross-functional teams and communicate technical concepts clearly to non-technical stakeholders.
Required Skills & Experience:
- 16 years of education with minimum 5+ years of hands-on experience in data engineering with cloud platforms (AWS preferred).
- Strong expertise in Databricks, Delta Lake, Apache Spark, and distributed data processing.
- Experience with Python, SQL, and ETL/ELT frameworks.
- Proven experience with data migration from legacy systems to cloud data lakes.
- Deep understanding of data modeling, curation layers, and consumption patterns.
- Familiarity with ML workflows, feature engineering, and model operationalization.
- Experience with DevOps, Git, CI/CD, and job orchestration tools.
- Excellent communication skills with the ability to translate complex concepts into clear business language.
Preferred Qualifications:
- Experience with AWS Databricks, Azure Data Factory Glue, Airflow, Kafka, Informatica, or similar ingestion tools.
- Knowledge of Unity Catalog, Delta Sharing, and enterprise governance frameworks.
- Exposure to AI/ML platforms, MLOps, or Databricks Feature Store.
- Certifications: Databricks Data Engineer Associate/Professional, AWS and Azure Data Engineer.