Oracle is building and expanding its next-generation Platform as a Service cloud offering. As a Principal Site Reliability Engineer, you will help operate, support, and improve Oracle Exadata Cloud Service, while leading the resolution of critical production issues and influencing the architecture of new service capabilities.
Responsibilities:
- Design, develop, test, and deliver software and automation that improve the availability, scalability, latency, security, operability, and efficiency of Oracle Database as a Service offering
- Lead the investigation and resolution of complex technical issues spanning Exadata Cloud Service, Autonomous Database, Oracle Database, operating systems, virtualization, storage, networking, and cloud infrastructure
- Coordinate response to high-severity incidents, including technical diagnosis, mitigation, stakeholder communication, recovery, and post-incident review
- Perform detailed root-cause analysis and develop corrective and preventive solutions that reduce the likelihood and impact of recurrence
- Build automation to eliminate repetitive operational work, reduce human error, accelerate incident response, and improve fleet-management efficiency
- Apply AI-assisted engineering and operations techniques to improve anomaly detection, incident correlation, troubleshooting, knowledge retrieval, capacity forecasting, and operational decision-making
- Evaluate and integrate generative AI, machine learning, and large language model capabilities into appropriate SRE workflows while maintaining security, privacy, accuracy, and human oversight
- Develop tools that use operational telemetry, logs, metrics, traces, events, and historical incident data to identify patterns and provide actionable insights
- Define, implement, and continuously improve service-level indicators, service-level objectives, error budgets, alerts, dashboards, and operational health metrics
- Improve monitoring and observability across distributed database and infrastructure services
- Participate in the architecture, design, implementation, and operational-readiness review of large-scale distributed DBaaS features
- Conduct research, prototyping, and proof-of-concept development for new service capabilities, reliability improvements, automation frameworks, and AI-enabled operational tools
- Act as a trusted technical advisor to customers and internal stakeholders, helping solve complex database, infrastructure, cloud, and DevOps challenges
- Create and deliver best-practice recommendations, sample code, runbooks, troubleshooting guides, technical documentation, and operational procedures
- Contribute to making Oracle’s cloud infrastructure simple, reliable, secure, and easy to operate
- Participate in capacity planning, demand forecasting, performance analysis, workload characterization, and system tuning
- Support system provisioning, patching, upgrades, maintenance, and lifecycle-management activities
- Review service and software designs for reliability, resilience, scalability, diagnosability, and operational sustainability
- Partner with development teams to improve product quality, supportability, automation, and customer experience
- Influence product roadmaps by translating operational findings and customer feedback into engineering requirements
- Mentor engineers and provide technical leadership during complex projects, incidents, and architectural discussions
- Establish and promote standard engineering practices, operational procedures, and reliability principles across the organization
- Participate in an on-call rotation and provide escalation support for critical production incidents
Requirements:
- Bachelor's degree in computer science, Computer Engineering, Information Systems, Management Information Systems, or another relevant technical field, or equivalent practical experience
- Typically, 8 or more years of experience in software engineering, site reliability engineering, systems engineering, database engineering, cloud operations, system administration, or a related technical discipline
- Advanced programming and scripting skills using Python, Shell, or Perl
- Ability to design, develop, troubleshoot, and maintain production-quality software and automation
- Strong Linux systems engineering experience and a detailed understanding of operating-system concepts, including processes, memory, storage, I/O, filesystems, networking, performance, and kernel behavior
- Experience building, operating, and troubleshooting virtualized or containerized systems using technologies such as KVM, Oracle VM, Docker, or similar platforms
- Ability to read, analyze, and troubleshoot source code across multiple components and technology layers
- Demonstrated experience diagnosing complex software, database, operating-system, storage, and networking issues
- Strong understanding of distributed systems, cloud-computing concepts, high-availability architectures, fault tolerance, scalability, and multi-tenant platforms
- Experience designing or operating highly available services governed by strict service-level objectives or service-level agreements
- Strong analytical and problem-solving skills, including the ability to isolate root causes across complex, interdependent systems
- Experience developing monitoring, observability, alerting, or operational telemetry solutions
- Experience automating production operations, incident response, system maintenance, or fleet-management activities
- Proven ability to learn unfamiliar technical domains quickly and transfer that knowledge to others
- Understanding of cloud networking concepts, including virtual cloud networks, security lists, network security groups, route tables, gateways, private connectivity, and network segmentation
- Knowledge of network security principles, including encryption in transit, TLS, certificate management, access controls, and secure network architecture
- Strong understanding of TCP/IP networking, routing, switching, DNS, DHCP, VLANs, subnets, firewalls, load balancers, proxies, and network address translation
- Experience troubleshooting complex network connectivity, latency, packet loss, throughput, and name-resolution issues across distributed cloud environments
- Strong verbal and written communication skills
- Ability to communicate complex technical findings clearly to engineering teams, leadership, customers, and other stakeholders
- Demonstrated technical leadership, sound judgment, and the ability to operate effectively during high-pressure production incidents
- Working knowledge of artificial intelligence, machine learning, generative AI, and large language model concepts
- Experience using AI-assisted development tools to improve coding, testing, debugging, documentation, or operational analysis
- Ability to identify practical and responsible uses of AI within site reliability engineering and cloud operations
- Understanding of retrieval-augmented generation, prompt design, model evaluation, embeddings, vector search, or AI agent workflows
- Ability to integrate AI or machine learning services through APIs, software development kits, or cloud-native services
- Experience analyzing operational data for anomaly detection, event correlation, failure prediction, capacity forecasting, or incident classification
- Understanding of the limitations and risks of AI-generated outputs, including hallucination, data leakage, prompt injection, model bias, non-determinism, and insufficient explainability
- Ability to design AI-assisted operational workflows with appropriate validation, auditability, access controls, security protections, and human approval mechanisms
- Commitment to using AI responsibly and in accordance with Oracle security, privacy, intellectual-property, and data-governance requirements
- Experience with Oracle Database technologies, including Oracle Real Application Clusters, Data Guard, Oracle Grid Infrastructure, Clusterware, Automatic Storage Management, and Recovery Manager
- Significant experience with Oracle Exadata or Exadata Cloud Service
- Experience supporting Autonomous Database or other Oracle Database as a Service offering
- Experience with Oracle Cloud Infrastructure services, APIs, software development kits, command-line interfaces, and operational tooling
- Experience with infrastructure-as-code and configuration-management technologies
- Experience with Kubernetes, containers, and cloud-native service architectures
- Programming experience with Java, C, or C++
- Experience working in cloud technical support, production engineering, network operations centers, service operations, or similar environments
- Experience developing or operating AI-enabled systems, machine learning pipelines, semantic-search platforms, copilots, or intelligent automation
- Knowledge of MLOps, model monitoring, AI observability, and model-lifecycle-management practices
- Experience building secure automation for regulated, mission-critical, or customer-facing environments
- Master's degree in computer science, Computer Engineering, Data Science, Artificial Intelligence, or a related discipline