Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. They are seeking a Senior Site Reliability Engineer to own high-impact infrastructure work, shape reliability and automation practices, and act as a technical bridge between development and operations.
Responsibilities:
- Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking
- Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on
- Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization
- Drive down cost and toil through better defaults, right-sizing, and automation rather than manual intervention
- Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do
- Define the reliability signals that matter for the services you own, and hold the line on them
- Develop and maintain disaster recovery and business-continuity strategies, and prove they work
- Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety
- Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows
- Build custom tooling and scripts that remove recurring operational pain
- Partner with data and development teams to improve deployment practices and application reliability
- Provide escalation support for production incidents, help lead post-incident reviews, and turn findings into durable fixes
- Participate in technical design reviews and offer architectural input across teams
- Help improve SRE and infrastructure best practices across the team, and participate in on-call for critical systems
Requirements:
- 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision
- GCP depth: strong hands-on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM). You know how these fit together in production, not just in a certification
- Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads
- Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code
- A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts
- Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one
- Distributed systems fundamentals: failure modes, consistency, and how systems break at scale
- Git and collaborative development workflows: you work in shared codebases and review others' changes well
- Incident management: you've run incidents and post-mortems and can stay calm and methodical when production is on fire
- Ownership & Accountability: You own features end-to-end and take pride in what you ship. You follow through from design to production and don't drop things
- Strong Communication: You can explain technical decisions and trade-offs to engineers, PMs, and stakeholders. You ask good questions and listen well
- Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates
- Production Mindset: You prioritize reliability and user impact. You think about failure modes, monitoring, and operational concerns as part of your design process
- Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools. You stay current without chasing hype
- Directed AI-Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality and your own technical judgment
- Database operations depth: PostgreSQL internals (logical replication, vacuuming, lock contention) or experience with migrations and database scaling. Familiarity with Redis, ClickHouse, or analytical stores is a plus
- Event-driven systems: Kafka/Redpanda or Pub/Sub, schema registries, and the operational realities of streaming at scale
- Cost engineering: you've meaningfully reduced cloud or observability spend without sacrificing reliability
- GCP certifications (Cloud Architect, Cloud DevOps Engineer) or demonstrably equivalent depth
- Cloudflare: experience with Workers and other Cloudflare services
- Multi-cloud or hybrid architecture exposure