Fortress Information Security is seeking a Senior Site Reliability Engineer to ensure the reliability, performance, and scalability of critical systems and services across AWS and on-premises environments. This role focuses on improving deployment processes, strengthening CI/CD pipelines, increasing automation, and enhancing system monitoring and observability.
Responsibilities:
- Transition applications from traditional (non-containerized) Ansible deployments to containerized, orchestrated deployments in AWS and on-premises environments
- Build upon current CI/CD efforts to support different deployment strategies (Blue/Green, Canary, etc.)
- Support Development/QA/UAT efforts by building an environment for anyone in the company to test drive a release anywhere in the development lifecycle
- Improve common infrastructure for developers, such as CI/CD pipelines, log/application monitoring, cluster management, and configuration management
- Migrate application secrets and configuration from Ansible Vault to Hashicorp Vault
- Automate infrastructure provisioning during deployment
- Handle code deployments in all environments (cloud, on-premises)
- Implement tools to monitor and alert with respect to service level metrics and objectives
- Provide technical guidance and educate team members and coworkers on development and operations
- Monitor relevant systems for availability and performance
- Available to support daytime and after business hours release activities
Requirements:
- 5-8 years hands-on experience in a SRE/DevOps role supporting production systems
- Production experience supporting Linux-based infrastructure and administering services on AWS (RDS, VPC, ECR, CloudWatch, Cloud Formation, Lambda, API Gateway) and on-premises
- Demonstrable experience with deployment technologies such as Kubernetes, Ansible, Jenkins and Terraform (or similar technologies)
- Excellent documentation skills so anyone on the team can come up-to-speed on changes quickly
- Excellent written and verbal communication skills
- Experience implementing rolling upgrades (canary, blue/green, etc.)
- Strong scripting and tooling skillset (Bash, Python, JS, etc.)
- Self-motivated, resourceful and a persistent problem-solving aptitude with advanced time management skills
- Ability to independently use and refine prompts to enhance the quality, efficiency, and insight of regular work processes
- Must be willing to participate in technical interviews and technical questions, which may be recorded or transcribed for evaluation purposes
- Bachelor's Degree in Information Technology, Computer Science, or a related discipline required
- SOC, NERC, and other compliance standards knowledge a plus
- Presentation skills and the ability to communicate with a variety of technical and non-technical audiences
- Strong Linux system administration skills
- Adept at creating comprehensive documentation
- AWS administration