Upstart is a leading AI lending marketplace focused on reducing the cost and complexity of borrowing for Americans. The Senior Engineering Manager of Site Reliability Engineering will lead a team to enhance the reliability and operational maturity of Upstart’s products and services, driving improvements in incident management, observability, and reliability engineering.
Responsibilities:
- Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
- Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
- Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
- Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
- Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
- Set a high bar for technical quality, operating rigor, and executive communication
- Develop engineers and leaders who can independently own complex reliability initiatives
- Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery
- Establish clear standards for managing high severity incidents and provide visible leadership during critical events
- Improve postmortem quality and ensure incident learnings result in durable engineering improvements
- Identify recurring failure patterns and drive systemic solutions across teams
- Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction
- Improve the quality, accessibility, and trustworthiness of signals used to understand production health
- Drive consistent practices across metrics, logs, traces, alerting, and service health
- Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions
- Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil
- Define measurable reliability outcomes and use data to prioritize investments and communicate impact
- Partner with platform and product engineering teams to embed reliability into standard engineering workflows
- Establish scalable operational readiness standards for new services, major launches, and architectural changes
- Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response
- Identify systemic reliability risks and partner with engineering teams to prioritize and address them
- Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards
- Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through
- Align stakeholders and dependencies before critical launches and engineering decisions
Requirements:
- 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
- Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
- Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
- Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
- Experience leading high severity incident response and improving incident management practices at scale
- Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
- Track record of hiring, developing, and retaining high performing engineers and engineering leaders
- Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams
- Experience operating large scale, highly available distributed systems
- Experience implementing or evolving service-level objectives and error-budget practices
- Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
- Experience developing incident management, operational readiness, or resilience programs across a large engineering organization
- Familiarity with Kubernetes, AWS, and modern cloud native architectures
- Experience supporting major platform or architectural transitions
- Strong product mindset when building internal reliability capabilities
- Experience establishing executive level reliability reporting and operating reviews