Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. The Azure Event Grid engineering team is seeking a talented and highly motivated Principal Software Engineer - Architect to lead the design and implementation of the next generation of PubSub solutions for customers worldwide.
Responsibilities:
- Architects extensible and maintainable solutions that emphasize diagnosability, reliability, and resilience at scale
- Drives requirements and design by partnering with stakeholders, including program managers, technical leads, and architects, to define and refine messaging system requirements. Leverages telemetry, customer feedback, and usage patterns to inform architectural decisions and shape the product roadmap. Establishes continuous feedback loops to measure customer value, reliability, and operational health, ensuring future design iterations are data-driven
- Champions coding best practices, design patterns, and reusable frameworks across the team. Ensures production-ready code with minimal defects and mentors engineers through hands-on guidance and comprehensive code reviews
- Defines testing strategies for messaging system components, establishing quality gates and success criteria across unit, integration, and end-to-end testing. Drives improvements in test coverage, removes obsolete tests, and identifies gaps in the testing framework. Leads the integration of automation into CI/CD pipelines to continuously validate reliability and performance under realistic workloads
- Improves engineering productivity by identifying tooling gaps across the cloud messaging development lifecycle. Designs and builds internal tools, frameworks, and libraries that streamline development, debugging, and operational workflows. Evaluates and advocates for open-source solutions where appropriate. Mentors engineers on modern tooling practices and fosters a culture of continuous improvement in developer experience
- Leads incident response and operational excellence as an Incident Manager. Monitors messaging systems for degradation, downtime, and service interruptions, driving rapid root-cause analysis and resolution of complex distributed systems issues. Coordinates cross-functional response efforts, communicates status to stakeholders, ensures SLA compliance, authors post-incident reviews, and drives systemic improvements to prevent recurrence
Requirements:
- Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience
- Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter
- Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 15+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience
- Hands on experience of developing AI Agentic solutions
- Deep understanding in distributed systems, strong analytical skills, and a proven track record of delivering large-scale services end to end preferably Pub/Sub and messaging services