Forge is a global solutions provider seeking a Mid Data Engineer to support legacy-to-modern data transformation in a secure AWS environment. The role involves developing data pipelines, automating data quality, and ensuring high-quality data flows across various interfaces.
Responsibilities:
- Build secure Python and AWS ETL/ELT pipelines for ingestion, transformation, reconciliation, and delivery
- Develop and evolve relational data models, schemas, indexes, constraints, views, and access patterns for MariaDB, PostgreSQL, or comparable platforms
- Develop data workflows and interfaces using Python on AWS Lambda and PySpark for event-driven, batch, and distributed transformation workloads
- Create automated data-quality checks for accuracy, completeness, consistency, timeliness, uniqueness, and business-rule conformance
- Implement source-to-target mapping, lineage, auditability, restartability, exception handling, and controlled replay
- Develop parity tests that compare legacy and modern processing outcomes and document the disposition of intentional differences
- Tune SQL and pipeline performance for high-volume batch and near-real-time workloads while protecting transactional integrity
- Implement monitoring, logging, alerting, and operational dashboards for pipeline health, latency, failures, and data quality
- Automate CI/CD, version-controlled data changes, deployments, rollback, and operational recovery controls
- Collaborate with architects, mission SMEs, Appian developers, testers, security personnel, and interface partners
Requirements:
- U.S. Citizen (Authorization to Work in the U.S. will not suffice); previous professional experience supporting the U.S. Federal Government, either as a federal employee or contractor, is required
- 4+ years of professional experience in data engineering, database development, or data-platform delivery
- Bachelor's degree in Computer Science, Information Systems, Data Engineering, or equivalent, OR 4 additional years of relevant professional experience in lieu of a degree
- Advanced Python software-engineering skills and experience building AWS Lambda functions and PySpark data-transformation pipelines
- Experience building, testing, and operating production ETL/ELT pipelines with automated data-quality controls
- Experience with data modeling, schema migration, source-to-target mapping, lineage, reconciliation, and data-quality automation
- Experience integrating data platforms with REST APIs, application services, file exchanges, and event-driven interfaces
- Experience with Git, CI/CD, automated testing, logging, monitoring, performance tuning, and production support
- Active CompTIA Security+ or equivalent DoW-approved baseline cybersecurity certification, or ability to obtain within the first 30 days of starting
- Active Tier 2 background investigation or higher, completed or favorably adjudicated within the previous 18 months
- The Candidate must be able to think strategically, act tactically, and demonstrate strong analytical and critical-thinking skills
- They must also build strong cross-group working relationships and demonstrate exceptional organizational skills and attention to detail
- The Candidate must be able to thrive and succeed in an entrepreneurial environment and not be hindered by ambiguity or competing priorities
- We are looking for self-managing candidates who enjoy working collaboratively in a fast-paced environment and working with dynamic teams
- Experience using Palantir Foundry for data integration, transformation, lineage, governance, and operational workflows
- Active Secret security clearance preferred
- Previous professional experience supporting a DoW organization, mission, or customer is strongly preferred; experience in modernizing COBOL flat files or legacy relational data into a modern relational architecture
- Experience integrating with Appian, Python microservices, financial transactions, logistics workflows, or high-volume external interfaces
- Experience serving as a technical team lead, mentoring junior engineers, or assisting teammates across delivery tasks
- Experience with BI, analytics, archival, records-retention, or NARA-aligned data lifecycle requirements