IT Applications Solutions Architect Mid / SRE Engineer
Location: Remote with Quarterly PI Planning Travel
Duration: 12 Months+
Client is looking for a DevOps/Site Reliability Engineer (SRE) with strong AWS cloud, automation, and production support experience. This is a hands-on engineering role.The ideal candidate will have experience building CI/CD pipelines (GitHub Actions, Jenkins, AWS CodePipeline), Infrastructure as Code (Terraform, CloudFormation, or AWS CDK), scripting with Python, and supporting cloud infrastructure in AWS (Azure exposure is a plus). Candidates should also have experience with Dynatrace observability, incident response, root cause analysis, ServiceNow, Linux, containers (Docker/Kubernetes/ECS), and SRE concepts such as SLIs, SLOs, error budgets, resiliency, and performance monitoring. Strong troubleshooting skills, automation experience, and the ability to support production environments and participate in on-call rotations are key to success.
Labor Category / Position Title
IT Applications Solutions Architect Mid / SRE Engineer
Skill Set
- Deployment & Automation
- Implement CI/CD pipelines using tools such as GitHub Actions, AWS Code Pipeline, and Jenkins
- Automate infrastructure provisioning through Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Develop automation scripts and self-service tools to enhance operational efficiency. - Leveraging Dynatrace Observability Platform
Demonstrated expertise in
- Standardized installation through automation
- Integrating with CI/CD pipelines
- Enforcing Tagging and Metadata standards
- Use of environment-aware configuration
- Implement distributed tracing with appropriate context propagation.
- Optimize alerts, create dashboards, alerts, and anomaly detectors - Incident Management & Response
- Proficient in ITIL framework and ITSM tools such as ServiceNow.
- Production on-call responder with strong troubleshooting capabilities.
- Develop RCA documentation, and Knowledge articles
- Apply SRE principles, including SLIs, SLOs, and error budgets. - Capacity Planning & Performance
- Implement operational cost optimization initiatives.
- Configure and maintain auto-scaling policies and thresholds.
- Develop Resiliency Test plans and support Performance testing. - Security & Compliance Implementation
- Manage service accounts and access permissions
- Create, deploy, and manage digital certificates.
- Respond to security incidents and execute remediation tasks effectively. - Education & Experience
- Bachelor s degree in computer science, Engineering, or related field
- 2 to 4 years of experience in DevOps, SRE, or infrastructure roles
- Mid-level proficiency in Python or other scripting languages.
- Mid-level proficiency in Configuration management tool including Ansible.
- Practical experience with cloud platforms AWS and Azure.
- Knowledge of containerization (Docker, Kubernetes/ECS).
- Knowledge of Linux systems and networking.
- Knowledge of relational, cloud, and NoSQL databases.
- Excellent written and verbal communication skills.
- Demonstrated ability to work independently and manage priorities.
- Availability to work outside of standard business hours as required.