
Reliability Engineering & Operations
Ensure high availability and reliability of enterprise applications in a 247 production environment.
Monitor applications, batch jobs, and workflows to maintain operational continuity.
Incident, Problem & Change Management
Lead and manage major incidents (P1/P2) and drive resolution to minimize business impact.
Perform root cause analysis (RCA) and implement preventive measures.
Ensure adherence to SLA/SLO and ITIL-based incident, problem, and change management processes.
Monitoring & Observability
Design and maintain monitoring dashboards.
Implement proactive alerting and improve system observability.
Troubleshooting & Support
Diagnose and resolve application and data-related issues using SQL queries and log analysis.
Provide backend validation and technical support across distributed environments.
Release & Deployment Support
Support release deployments, change validation, and post-deployment activities.
Participate in disaster recovery testing and release readiness validation.