ampliFI Loyalty Solutions is a company focused on providing customized credit and debit card loyalty programs for banks and credit unions. They are seeking a Site Reliability Engineer to bridge the gap between software development and IT operations, focusing on optimizing the release lifecycle and ensuring the stability and scalability of AWS-based cloud infrastructure.
Responsibilities:
- Lead and coordinate the end-to-end release lifecycle, including planning, scheduling, staging, deploying, and post-release validation
- Act as a key representative and technical coordinator on the Change Advisory Board, defending upcoming releases, evaluating architectural risk, and ensuring all compliance requirements are met prior to production deployment
- Serve as the primary point of contact for developers, QA, product management, and business stakeholders regarding deployment windows, release status, risk assessments, and rollback plans
- Standardize and mature release processes, transitioning manual gatekeeping into automated CI/CD guardrails and repeatable workflows
- Provide hands-on tier-2 production support, ensuring operational stability and participating in the active engineering on-call rotation
- Design, build, and maintain internal scripts, custom tooling, and automated pipelines (Python, Bash, PowerShell) to reduce operational 'toil' and streamline release operations
- Respond promptly to system outages, service interruptions, and security alerts. Lead post-incident Root Cause Analysis (RCA) efforts and implement permanent preventive engineering solutions
- Maintain, deploy, and scale AWS cloud infrastructure using Terraform, AWS CloudFormation, or Ansible to ensure environment parity and drift-free deployments
- Configure, tune, and optimize monitoring and alerting systems (e.g., CloudWatch, Prometheus, Grafana) to provide comprehensive visibility into release health and production performance
- Ensure all deployment and release activities strictly adhere to PCI, SOC2, and internal corporate security/governance standards
- Work closely with software engineering, QA, and platform infrastructure teams to build 'paved paths' for developers to ship software safely and rapidly
- Maintain pristine, audit-ready documentation for release logs, standard operating procedures (SOPs), runbooks, and incident timelines
Requirements:
- 3–5 years of experience in SRE, DevOps, system administration, or Release/Operations engineering roles
- Proven experience coordinating software releases, managing multi-tier deployment pipelines, and working within formal ITIL/Change Management frameworks (including active CAB participation)
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent practical experience)
- Hands-on experience configuring, deploying, and maintaining AWS infrastructure and serverless architectures (EC2, S3, RDS, IAM, Lambda, API Gateway)
- Solid understanding and working experience with IaC tools such as Terraform or AWS CloudFormation
- Strong proficiency in scripting languages (especially Python and Bash) to write automation tooling and integrate systems
- Demonstrated experience leading Root Cause Analysis (RCA) and participating in production on-call rotations
- Exceptional written and verbal communication skills, with a proven ability to coordinate across highly technical development teams and business-oriented leaders
- AWS Certified SysOps Administrator, AWS Certified DevOps Engineer, or ITIL Foundation certifications
- Experience configuring and maintaining CI/CD systems (such as GitLab CI, GitHub Actions, Jenkins, or AWS CodePipeline)
- Hands-on experience with modern monitoring, APM, and alerting tools (e.g., Datadog, Prometheus, Grafana, PagerDuty)
- Experience with Docker and container orchestration platforms (Kubernetes, AWS ECS/EKS)
- Hands-on experience assisting with SOC2 or PCI-DSS audits, especially documenting evidence for code changes and release gates