Akraya, Inc. is an award-winning IT staffing firm recognized for excellence in the industry. They are seeking a Senior Staff Engineer – NPE Observability to lead the architecture and evolution of a large-scale telemetry and observability platform, focusing on real-time data ingestion, streaming, and visualization across global network infrastructure.
Responsibilities:
- Architect and optimize scalable telemetry ingestion and storage platforms using Java and PostgreSQL
- Design and enhance high-throughput Apache Kafka streaming pipelines for real-time telemetry processing
- Define enterprise observability standards and build advanced Grafana dashboards for monitoring global infrastructure
- Architect stateful stream-processing solutions using technologies such as Apache Flink
- Evaluate and prototype emerging observability technologies including Model-Driven Telemetry (MDT), ClickHouse, and Thanos
- Define platform architecture, technical standards, and long-term engineering roadmaps
- Establish and monitor SLIs/SLOs to ensure high availability, reliability, and platform performance
- Lead complex root cause analysis, performance tuning, and architectural improvements for mission-critical systems
- Collaborate with software engineering, network engineering, and infrastructure teams to translate business requirements into scalable technical solutions
- Mentor engineering teams and drive technical excellence through architecture reviews and best practices
Requirements:
- 10+ years of software engineering experience with expertise in distributed systems
- 5+ years of experience building large-scale network engineering, telemetry, or observability platforms
- Expert-level proficiency in Java backend development
- Strong expertise with Apache Kafka, including cluster architecture, messaging, and stream processing
- Advanced experience with PostgreSQL schema design, optimization, and performance tuning
- Expert-level experience developing enterprise dashboards using Grafana
- Strong understanding of distributed systems, real-time streaming, and high-throughput data processing
- Experience with Prometheus, Thanos, ClickHouse, or similar observability platforms
- Experience defining SLIs, SLOs, monitoring strategies, and incident management
- Strong stakeholder management, technical leadership, and architectural design skills
- Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical discipline
- Experience designing globally distributed, high-availability observability platforms
- Strong background in Agile/Scrum software development methodologies
- Proven ability to drive long-term technical vision, innovation, and engineering excellence in enterprise-scale environments