Cisco is revolutionizing how data and infrastructure connect and protect organizations in the AI era. The Site Reliability Engineer role involves ensuring the reliability and scalability of network and cloud infrastructure while supporting production services and improving incident management processes.
Responsibilities:
- Works independently, but receives minimal guidance and direction from leader then determines best approach to accomplish work
- Acts as a resource for colleagues with less experience
- Understands project and/or department needs and establishes relationships with appropriate cross-functional stakeholders to gather input, collect information, and complete work steps
- Designs and deploys small to mid-size or moderately complex solutions to optimize reliability, availability, latency, and performance
- Integrates knowledge of design, automation, and deployment with expertise in coding to improve service reliability for existing or new systems and adapts for regions, countries, or customers
- Designs and tests high availability and disaster recovery measures for our services to ensure automation is improving reliability, scalability, and velocity
- Forecasts and builds reports to determine at what point resources will be at capacity
- Designs and implements tools that provide visibility into performance and reliability of our infrastructure
- Builds automated platforms
- Monitors the environment and works with Developers and Ops to identify problems and develop monitoring tools that provide visibility into performance and reliability, serves as on-call SRE, leads post mortems, and writes root cause analysis
- Builds and ensures security controls are in place in regards to architectural design, collaborates with security in designing or providing input to security controls, and may actively contribute in security incident response
Requirements:
- 7+ yrs experienced professional using best practices and knowledge of internal or external business issues to improve products or services
- Works independently, but receives minimal guidance and direction from leader then determines best approach to accomplish work
- Acts as a resource for colleagues with less experience
- Understands project and/or department needs and establishes relationships with appropriate cross-functional stakeholders to gather input, collect information, and complete work steps
- Designs and deploys small to mid-size or moderately complex solutions to optimize reliability, availability, latency, and performance
- Integrates knowledge of design, automation, and deployment with expertise in coding to improve service reliability for existing or new systems and adapts for regions, countries, or customers
- Designs and tests high availability and disaster recovery measures for our services to ensure automation is improving reliability, scalability, and velocity
- Forecasts and builds reports to determine at what point resources will be at capacity
- Designs and implements tools that provide visibility into performance and reliability of our infrastructure
- Builds automated platforms
- Monitors the environment and works with Developers and Ops to identify problems and develop monitoring tools that provide visibility into performance and reliability, serves as on-call SRE, leads post mortems, and writes root cause analysis
- Builds and ensures security controls are in place in regards to architectural design, collaborates with security in designing or providing input to security controls, and may actively contribute in security incident response
- Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PhD + 1 year of related experience
- Requires solid conceptual and practical knowledge in primary technical job family and knowledge of related technical job families; has worked with a range of technologies