TSPi is seeking a Data Engineer to design, develop, and maintain scalable data pipelines and analytics workflows. The role involves leveraging cloud-based data analytics platforms to build secure data solutions that support analytics and operational needs.
Responsibilities:
- Design and implement scalable data pipelines for ingesting, transforming, and processing large datasets
- Develop ETL/ELT workflows supporting enterprise data integration and analytics initiatives
- Build and maintain data processing workflows using Databricks and distributed computing technologies
- Develop data transformation and analytics solutions using Python, PySpark, and SQL
- Ensure data quality, consistency, integrity, and governance across systems and workflows
- Optimize performance of data pipelines and large-scale data processing workloads
- Collaborate with analytics, engineering, DevOps, and business teams to support data platform operations and modernization efforts
- Maintain documentation for data architecture, data models, and pipeline processes
- Support secure handling and management of sensitive or regulated data
- Contribute to automation and AI-enabled data workflow initiatives where applicable
Requirements:
- Experience working with cloud-based data analytics platforms, specifically Databricks or similar technologies
- Strong programming experience in: Python
- Strong programming experience in: PySpark
- Strong programming experience in: SQL
- Experience building distributed data processing pipelines
- Strong understanding of ETL/ELT pipeline development and data transformation processes
- Experience working with large and complex datasets
- Experience supporting cloud-native or modernized data environments
- Strong analytical, troubleshooting, and problem-solving skills
- Experience collaborating within cross-functional Agile teams
- Excellent written and verbal communication skills
- US work authorization required and ability to obtain and maintain a Public Trust clearance, which may include fingerprinting
- Experience implementing data lake or lakehouse architectures
- Experience working within AWS, Azure, or Google Cloud Platform environments
- Experience with workflow orchestration and automation tools
- Experience supporting federal, healthcare, public sector, or highly regulated environments preferred
- Familiarity with AI/ML data workflows or large language model integrations is a plus