Fundraise Up is a global fundraising platform focused on making donations to nonprofits fast and accessible. The Senior DevOps Engineer will be responsible for owning core platform areas, driving technical initiatives, and mentoring less experienced engineers while ensuring the reliability and scalability of the systems.
Responsibilities:
- Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap
- Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards
- Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps
- Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model
- Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels
- Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand
- Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works
- Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage
- Use AI in all aspects of day-to-day work: researching, troubleshooting, developing
Requirements:
- 6+ years as a DevOps Engineer / SRE (or very close responsibilities)
- Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery. You can showcase initiatives that were yours, not just tasks you completed
- Confident Linux skills (we use Ubuntu)
- Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup
- Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery
- Containers: Docker, image building, registries
- Ansible
- Git
- Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work)
- Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems
- Experience mentoring less experienced engineers
- Ownership and attention to detail. Downtime is expensive: during busy events 10 minutes of downtime can cost us around $500k
- Must Have: solid hands-on experience in two or more of the areas below: VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure
- Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance
- Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets
- Container registries and artifact management: Harbor, Nexus, base images, image policies
- Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting
- Grafana: dashboards as code, alerting, performance at scale
- Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits. Building platform around these tools to improve Quality of Life for Analytics
- Remote development environments and AI agent execution environments. E.g. Coder/Telepresence
- Bare-metal Kubernetes: provisioning, networking, scaling
- Flux and GitOps
- Terraform
- Sentry on-premise: operating self-hosted error tracking
- ClickHouse, MongoDB