Arkestro is a hyper-growth company that applies AI, game theory, and behavioral science to enterprise negotiations. They are seeking a Senior Site Reliability Engineer to manage performance and reliability for their software platform and infrastructure, collaborating with various teams to improve and maintain their systems.
Responsibilities:
- Building and setting up new development tools and infrastructure
- Working on ways to automate and improve development and release processes
- Management of existing cloud infrastructure
- Document and act as subject matter expert for practices and policies involving infrastructure
- Improve reliability, quality, and time-to-market of our suite of software solutions
- Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement
- Provide primary operational support and engineering for distributed software applications
- Continuously integrate production context back into the development lifecycle, ensuring the system learns from past incidents to prevent recurring issues
Requirements:
- 5+ years of experience with various AWS products, like S3, RDS, Elastic Containers, etc
- 5+ years of experience with Kubernetes or other orchestration products
- 3+ years of experience with Infrastructure as Code
- 3+ years of experience with using observability platforms and log management, like Datadog, Splunk, Rigor, etc
- 6 months+ of experience working with LLMs, prompt engineering, harness engineering, or AIOps tooling
- 6 months+ of SDLC AI native usage via Claude Code, Claude Cowork or similar
- Excellent communications skills, the ability to learn on the fly, and a desire for ownership