Job Title
Senior Kafka Engineer / Administrator (Onsite)
Location
Denver, CO
Duration
12 Months
Interview Type
Virtual & In-Person
Note
Must be local to Denver, CO or St. Louis, MO - local only 4days/wk in office
Description
About the Role
We are migrating a portfolio of Apache Kafka clusters from on-premises platforms (Cloudera and Confluent) to AWS MSK, managed through a CI/CD pipeline built on Terraform and GitLab. We are looking for a senior, hands-on Kafka engineer to own day-to-day Kafka administration across both environments, drive the migration, and coordinate closely with producer and consumer application teams to ensure a smooth, low-risk cutover.
This is a senior individual-contributor role for someone who has run Kafka at scale in production, is comfortable in a hybrid on-prem/cloud world, and can act as the technical point of contact between platform and application teams.
Required Qualifications
5+ years in data/platform/infrastructure engineering, with 4+ years operating Apache Kafka in production.
Hands-on administration experience with Cloudera and Confluent Platform.
Production experience with AWS MSK (provisioning, configuration, IAM auth, networking).
Proven experience migrating Kafka workloads to the cloud with minimal disruption.
Strong Terraform and GitLab CI/CD experience for infrastructure automation.
Hands-on Datadog experience for monitoring, dashboards, and alerting.
Deep understanding of Kafka internals: partitions, replication, ISR, offsets, consumer groups, and delivery semantics.
Solid AWS fundamentals: VPC, security groups, IAM, KMS, CloudWatch, PrivateLink.
Strong Linux, networking, and scripting skills (Bash, Python).
Excellent communication and coordination skills to work across producer/consumer teams.
Preferred / Nice to Have
Experience with MirrorMaker 2, Confluent Replicator, or cluster linking for replication/migration.
Schema Registry, Kafka Connect, and ksqlDB experience.
Kafka Streams or Flink familiarity.
Kubernetes / EKS experience (e.g., Strimzi).
Certifications: Confluent Certified Administrator, AWS Solutions Architect/DevOps.
Experience in a regulated or large enterprise environment.
Top Skills
Kafka Administration (On-Prem + Cloud)
Administer and operate Kafka clusters across Cloudera, Confluent Platform, and AWS MSK.
Manage topics, partitions, replication, retention, quotas, ACLs, and consumer groups.
Tune brokers, producers, and consumers for throughput, latency, and reliability.
Handle patching, upgrades, capacity planning, and cluster health/performance troubleshooting.
Configure and maintain security: TLS/SSL, SASL, mTLS, IAM (for MSK), RBAC/ACLs, and encryption at rest and in transit.
Migration to AWS MSK
Plan and execute migration of clusters and workloads from on-prem (Cloudera/Confluent) to AWS MSK.
Design cutover strategies (e.g., MirrorMaker 2 / Confluent Replicator / cluster linking) with minimal downtime and no data loss.
Validate topic parity, offsets, throughput, and data integrity pre- and post-migration.
Define rollback plans and run migration dry-runs in lower environments.
CI/CD & Infrastructure as Code
Build and maintain Kafka/MSK infrastructure using Terraform.
Manage automated deployments and configuration through GitLab CI/CD pipelines.
Implement GitOps practices for topic, ACL, and cluster configuration management.
Enforce version control, code review, and environment promotion (dev ? test ? prod).
Observability & Monitoring
Build and maintain monitoring, alerting, and dashboards in Datadog for Kafka/MSK (broker metrics, consumer lag, partition health, JVM, throughput).
Establish SLAs/SLOs and proactive alerting to reduce incidents and mean-time-to-resolution.
Correlate Kafka metrics with application and infrastructure telemetry.
Producer / Consumer Coordination
Serve as the primary technical liaison between the Kafka platform team and producer/consumer application teams.
Advise application teams on client configuration, serialization/schema management, partitioning strategy, idempotence, and exactly-once/at-least-once semantics.
Support onboarding of new producers/consumers and troubleshoot client-side issues (rebalancing, lag, poison messages).
Coordinate migration schedules, testing, and cutover windows with each impacted team.
Recruiter
Contact: Sohail -