Cisco is revolutionizing how data and infrastructure connect and protect organizations in the AI era. The Site Reliability Engineer is responsible for ensuring the reliability and security of production environments while supporting uptime through monitoring and automation.
Responsibilities:
- Evaluating, improving, and maintaining the reliability, scalability, and security of production environments
- Monitoring, incident management, disaster recovery, and automation to ensure service uptime and compliance with SLAs
- Crafting the network architecture for a hybrid cloud design
- Improving agility by automating endpoint network connectivity
- Creating the best connectivity solution by abstracting network functions from hardware
- Developing comprehensive monitoring tools that provide transparency into the performance and reliability of our network infrastructure
- Full-stack troubleshooting of production issues (application, system and network) through to root cause analysis and implementation of preventative measures
- Design, implementation and management of an overlay network to support 1000’s of containers
Requirements:
- Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PhD + 1 year of related experience
- Experience in designing, deploying and operating mid to large scale enterprise or cloud environments
- Comfortable with scripting or coding with languages like Python, Bash, Ruby, Go
- Knowledge of *nix systems and interests that span beyond routers and switches
- Ability to jump into other people's solutions and source code to seek a problem
- Passion for automation
- Care and empathize with the customer experience and can support an externally facing production environment
- Passionate about building complex solutions working alongside other engineering teams
- Empathize with coworkers and have a positive influence on others
- Comfortable using AI-assisted engineering tools such as Codex, Claude Code, and Cursor to support coding, code reviews, troubleshooting, testing, documentation, and incident investigation
- Experience with: BGP, OSPF, IPv6, Network Security, DMVPN, MACSec, Grafana, Splunk, advanced Unix/Linux, system administration, scripting/Bash/Ruby/Python, project management experience, AWS/Azure, Docker, K8s, Ansible, REST APIs, TLS or Cloud/ISP/Telco exposure
- Understand agent-based workflows, including reusable skills and MCP integrations, and can critically evaluate AI-generated output for accuracy, security, reliability, and maintainability
- Keywords: Network Engineering, Production Engineering, SRE, Site Reliability Engineering, DevOps, CCNP, CCIE, JNCP, JNCIE