Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their apps and websites. They are seeking two Site Reliability / Production Engineers to respond to live incidents, assess customer impact, and improve system reliability.
Responsibilities:
- Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state
- Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code
- Use AI throughout triage and diagnosis while checking its conclusions against real evidence
- Choose and execute a proportionate mitigation, rollback, repair or bounded fix
- Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green
- Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication
- Join customer conversations occasionally when direct technical involvement is genuinely useful
- Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix
- Escalate with evidence, customer impact, actions already taken and the specific decision or help required
- Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait
- Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners
- Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes
- Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements
- Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve
- Create safe, supervised automation for common operational actions
- Work with product teams to close observability, rollback, runbook and supportability gaps
- Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner
- Make reliability and on-call performance easier for the company to understand and improve over time
Requirements:
- Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed
- Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded
- Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan
- AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions
- Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals
- Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work
- Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand
- Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it
- Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response
- Cloud platforms such as Azure or Cloudflare
- Distributed application and API diagnostics
- Databases, queues and background-processing systems
- Observability, alerting and incident-management platforms
- Infrastructure, deployment and release automation
- Application development and safe production debugging
- AI coding agents and workflow automation