GXO Logistics, Inc. is a leading provider of cutting-edge supply chain solutions to the most successful companies in the world. As a Senior DevOps AI Platform Engineer, you will establish and scale the enterprise DevOps operating model for GXO's Agentic AI Platform, building secure, automated, and scalable cloud platform capabilities that enable AI application delivery.
Responsibilities:
- Establish the enterprise DevOps operating model for GXO's Agentic AI Platform, including CI/CD standards, branching strategies, release governance, environment promotion, deployment approvals, and operational handoff practices
- Design, build, and manage secure, repeatable CI/CD pipelines supporting AI platform infrastructure, platform services, agents, MCP servers, LiteLLM, Agent Gateway integrations, model-serving components, and supporting services
- Engineer, deploy, and operate Kubernetes-based platform capabilities on Google Kubernetes Engine (GKE), including deployment standards, Helm or Kustomize, autoscaling, network policies, workload identity, secrets management, ingress/egress, observability, and production runbooks
- Own Terraform infrastructure delivery by developing reusable modules, managing state, enforcing pull request controls, implementing policy guardrails, maintaining environment parity, detecting configuration drift, and promoting infrastructure across development, test, staging, and production environments
- Partner with the Principal Cloud Engineer to implement Google Cloud Platform foundations while leading day-to-day DevOps enablement, release engineering, Kubernetes operations, pipeline reliability, and developer experience
- Collaborate with the Principal Cloud AI Platform Architect to translate enterprise architecture standards, reference architectures, and architectural decision records (ADRs) into automated build, test, deployment, and operational processes
- Enable Phase 2 platform capabilities, including GKE-based open-source model serving, vLLM or comparable inference runtimes, scalable deployment patterns, model tiering infrastructure, and cost-governed platform operations
- Implement enterprise DevSecOps controls in partnership with Information Security, including vulnerability scanning, dependency scanning, container image hardening, Binary Authorization (or equivalent), secrets management, audit logging, and secure deployment gates
- Create standardized 'paved road' developer workflows that enable engineers to provision environments, deploy AI agents, publish MCP services, test integrations, and promote code changes through approved automation
- Champion AI-assisted software engineering practices by enabling secure AI coding tools, automated testing, documentation generation, code review acceleration, pipeline diagnostics, and developer productivity improvements
- Build comprehensive observability across the platform through logs, metrics, traces, dashboards, alerts, SLOs, SLIs, deployment health monitoring, traceability, cost attribution, and operational readiness reporting
- Automate operational processes to reduce manual effort, improve incident response readiness, and maintain runbooks for releases, rollbacks, break-glass procedures, platform operations, and escalation processes
- Support secure integration between the AI platform and Snowflake-governed data access patterns through automated deployment, configuration, policy enforcement, and runtime observability
- Develop and maintain engineering documentation, including CI/CD standards, Terraform module guidance, Kubernetes operating procedures, release checklists, onboarding documentation, and operational runbooks
Requirements:
- Bachelor's degree in computer science, Engineering, Information Technology, Cloud Computing, or a related technical field; equivalent hands-on experience may be considered
- Google Cloud Professional DevOps Engineer certification required
- Minimum of 8 years of platform engineering, DevOps, Site Reliability Engineering (SRE), infrastructure engineering, cloud engineering, or software delivery engineering experience
- Minimum of 5 years of hands-on Google Cloud Platform experience supporting production environments
- Deep expertise with Google Kubernetes Engine (GKE), including Kubernetes operations, workload identity, networking, autoscaling, ingress/egress, Helm or Kustomize, and production troubleshooting
- Expert-level experience developing and managing Terraform infrastructure, including reusable modules, state management, CI/CD integration, policy-as-code, infrastructure promotion, and drift management
- Strong experience designing and maintaining secure CI/CD pipelines using Cloud Build, GitHub Actions, GitLab CI, Azure DevOps, Jenkins, or similar platforms
- Experience implementing GitOps and DevSecOps practices, including code review automation, dependency scanning, container security, secrets management, signed artifacts, deployment approvals, and security guardrails
- Experience supporting cloud-native AI, machine learning, analytics, developer platform, or data platform workloads on Kubernetes and Google Cloud
- Ability to collaborate effectively with principal architects, cloud engineers, Information Security, product teams, and software developers to translate architectural vision into production-ready solutions
- Strong operational mindset with experience supporting incident response, root cause analysis, observability, production support, SLOs/SLIs, release readiness, and continuous operational improvement
- Excellent technical communication skills with the ability to develop engineering documentation, operating procedures, automation standards, and developer guidance
- Ability to influence engineering teams across a global matrix organization while driving adoption of modern DevOps and platform engineering practices
- HashiCorp Terraform Associate certification
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) certification
- Google Cloud Professional Cloud Architect, Cloud Security Engineer, or Machine Learning Engineer certifications
- Experience with AI platform technologies including LiteLLM, Agent Gateway, MCP servers, Vertex AI, Gemini, model routing, vLLM, or open-source model serving frameworks
- Experience enabling AI-assisted software development through secure coding assistants, automated testing, documentation generation, and developer productivity tooling
- Strong understanding of secure enterprise AI platform operations, cloud-native architecture, and scalable infrastructure automation
- Experience driving platform standardization, operational excellence, and developer enablement across large engineering organizations
- Self-starter with the ability to quickly establish credibility, operate independently, and make an immediate impact on the reliability, security, scalability, and velocity of enterprise AI platforms