Revenium is focused on ensuring that financial decisions made by autonomous systems are attributable, governable, and profitable. They are seeking a Senior Site Reliability Engineer / Platform Engineer to manage the reliability, security, and operability of their production platform on AWS, driving incident response and maintaining compliance for enterprise customers.
Responsibilities:
- AWS infrastructure and Terraform
- Observability, alerting, and on-call experience
- Identity, access, and lifecycle automation
- Security and compliance posture
- Production reliability and incident response
- CI/CD pipeline ownership
- Database, messaging, and stateful service operations
- Backup, disaster recovery, and restore testing
- Networking, DNS, CDN, and edge controls
- Cost optimization and FinOps
- Developer tooling and platform engineering
Requirements:
- 10+ years of professional experience operating production cloud infrastructure for a SaaS product, with at least 5 years on AWS in a senior or principal capacity
- Expert-level proficiency in leveraging AI to automate SRE work
- Terraform — fluent
- AWS depth across our stack
- Production incident response experience
- Observability experience
- CI/CD ownership experience
- Linux and shell proficiency
- Containerization knowledge
- Networking fundamentals
- Security fundamentals
- Identity and SSO integration experience
- Compliance posture experience
- Strong communication skills
- Direct technical ownership of an ongoing compliance program
- Experience operating Apache Kafka in production
- Experience executing major-version PostgreSQL upgrades in production
- Experience operating OpenTelemetry pipelines
- Experience operating Spring Boot / JVM applications from an SRE perspective
- Experience with a runtime security or EDR tool
- AWS certifications
- Comfortable reading Kotlin or Java
- Experience with AI-driven automation in operational practices
- Experience with CI/CD pipeline migrations
- Experience with cloud cost management