Title: Databricks Data Engineer
Location: Bethesda, MD -- Onsite/Remote: Hybrid
But prefer someone who can be onsite as needed.
duration: 6 months.
Role Overview:
We are seeking a skilled Databricks Data Engineer to design, build, and maintain scalable data pipelines and data platforms using Databricks, Apache Spark, Delta Lake, and cloud technologies. The ideal candidate will have strong experience in data engineering, ETL/ELT development, large-scale data processing, and modern lakehouse architectures.
Key Responsibilities:
Design, develop, and maintain scalable ETL/ELT data pipelines using Databricks.
Develop data processing solutions using Apache Spark, PySpark, and SQL.
Build and manage Delta Lake tables and implement Medallion Architecture (Bronze, Silver, and Gold layers).
Ingest data from structured and unstructured sources, including databases, APIs, files, streaming platforms, and cloud storage.
Optimize Spark workloads, Databricks clusters, SQL queries, and data pipelines for performance and cost efficiency.
Implement data quality checks, validation rules, monitoring, logging, and error-handling frameworks.
Develop reusable data engineering frameworks and utilities.
Work with Databricks Workflows/Jobs to schedule and orchestrate pipelines.
Implement data governance, security, access controls, and data lineage using Unity Catalog.
Collaborate with data architects, analysts, data scientists, application teams, and business stakeholders.
Support CI/CD processes for Databricks notebooks, code, and data pipelines.
Troubleshoot production data issues and ensure reliability, scalability, and availability of data solutions.
Participate in code reviews, technical design discussions, and data architecture decisions.
Document data pipelines, transformation logic, architecture, and operational procedures.
Required Skills:
Strong hands-on experience with Databricks.
Strong knowledge of Apache Spark and PySpark.
Advanced SQL and data transformation skills.
Experience with Delta Lake and Lakehouse architecture.
Experience designing and implementing ETL/ELT pipelines.
Strong understanding of data warehousing, dimensional modeling, and data lake concepts.
Experience working with Microsoft Azure
Experience with cloud storage technology - ADLS
Knowledge of Databricks Jobs, Workflows, notebooks, clusters, and compute configurations.
Experience with Unity Catalog, data governance, and role-based access control.
Familiarity with Git and CI/CD practices.
Strong troubleshooting, performance optimization, and problem-solving skills.
Preferred Skills:
Experience with Databricks Auto Loader and Structured Streaming.
Experience with Delta Live Tables / Lakeflow Declarative Pipelines.
Knowledge of Kafka or other streaming technologies.
Experience with orchestration tools such as Azure Data Factory, Airflow, or equivalent.
Experience with Terraform or other Infrastructure-as-Code technologies.
Knowledge of Python software engineering best practices.
Experience implementing data quality frameworks.
Familiarity with Databricks Asset Bundles and automated deployment practices.
Databricks certification is a plus.