Conexess Group is seeking a Senior Data Engineer to join its enterprise data platform team. The role involves designing and optimizing data pipelines, managing data product delivery, and ensuring data quality and governance across various AWS services.
Responsibilities:
- Design, build and optimize data pipelines using AWS Glue, PySpark and Apache Iceberg
- Own data product delivery from source-system ingestion through transformation, quality validation and governed consumption
- Build and maintain Amazon Redshift integrations, including schemas, stored procedures, materialized views and Liquibase-managed DDL migrations
- Write and optimize complex SQL for analytics transformations, reporting views and data-quality checks
- Translate operational data models into star schemas and other dimensional models for analytics
- Manage Apache Iceberg tables, including partition and schema evolution, compaction and orphan-file cleanup
- Implement table- and column-level governance using Lake Formation permissions, tag-based access control and PII classification
- Develop data-quality frameworks covering referential integrity, row-count reconciliation, completeness and anomaly detection
- Partner with data product teams to onboard new data sources, define schemas and establish data contracts
- Troubleshoot AWS Glue jobs, DMS replication issues and data-freshness problems
- Own the complete lifecycle of changes, including development, deployment, validation and communication
- Ensure pipelines and datasets remain accurate, reliable and healthy across environments
Requirements:
- At least five years of hands-on data engineering experience building ETL/ELT pipelines and analytics platforms
- Expert-level SQL skills, including complex joins, CTEs, window functions, query optimization and performance tuning
- Strong Python and PySpark development experience
- Hands-on AWS Glue experience, including job development, bookmarks, crawlers and performance optimization
- Strong Amazon Redshift experience, including schema design, stored procedures, Spectrum, external schemas, materialized views and performance tuning
- Experience with Apache Iceberg or another open table format, including ACID transactions, schema evolution, partition evolution and table maintenance
- Strong understanding of dimensional data modeling, including star schemas, snowflake schemas and slowly changing dimensions
- Experience with AWS data services such as Athena, Lake Formation, S3, DMS, Lambda and EventBridge
- Understanding of data governance, tag-based access control, data classification and PII handling
- Experience with Liquibase or a similar database migration tool
- Familiarity with GitHub and CI/CD pipelines
- Strong attention to data correctness, completeness and freshness
- Ability to independently deliver production-ready data pipelines with limited oversight
- Data Mesh architecture and federated governance
- Lakehouse architecture patterns
- AWS DMS and change-data-capture ingestion
- Terraform or other infrastructure-as-code tools
- Data catalog and lineage platforms such as AWS DataZone or IDERA ER/Studio
- AI-assisted development tools such as Amazon Kiro or GitHub Copilot
- Amazon SageMaker or other machine-learning platforms
- Docker and container-based development
- Agriculture, retail or other large-scale operational data environments