We are seeking an AI Systems Engineer to design, build, and operate production-grade agentic AI systems. The successful candidate will develop multi-agent orchestration workflows, retrieval-augmented generation (RAG) pipelines, and natural-language interfaces to enterprise data sources, and will own these systems through deployment, monitoring, and ongoing optimization.
This is a hands-on engineering role requiring equal strength in applied AI development and distributed systems engineering. The candidate will work closely with data engineering, security, and product teams to deliver AI capabilities that meet enterprise standards for reliability, security, observability, and cost efficiency.
- Design, develop, and deploy agentic AI systems in Python using LangGraph and LangChain, including supervisor and sub-agent architectures, tool routing, and human-in-the-loop workflows.
- Implement agent memory and context management strategies, including conversational state, long-term semantic memory, summarization, and context-window optimization.
- Integrate AI agents with internal and third-party systems using the Model Context Protocol (MCP) and Agent-to-Agent (A2A) protocols.
- Build and optimize RAG pipelines end to end, covering document ingestion, chunking, embedding, hybrid and semantic retrieval, re-ranking, and access-control-aware filtering.
- Develop text-to-SQL and natural-language analytics capabilities over large relational schemas, including semantic catalogs, query validation, and execution guardrails.
- Design and implement scalable backend services and microservices in FastAPI, including RESTful API design, request validation with Pydantic, dependency injection, authentication and authorization, API versioning, and rate limiting.
- Build real-time streaming interfaces using WebSockets and Server-Sent Events to support token-level LLM response streaming and long-running agent executions.
- Model, provision, and optimize application data stores, including relational databases such as PostgreSQL, vector databases for embedding storage and similarity search, document stores such as MongoDB, and caching layers such as Redis.
- Own database lifecycle work including schema migrations, ORM usage, connection pooling, transaction management, and query performance tuning under production load.
- Containerize applications with Docker and deploy to Kubernetes, managing autoscaling, resource allocation, secrets, and progressive rollout strategies.
- Implement and maintain LLM gateway and routing infrastructure, including multi-provider failover, rate limiting, budget enforcement, and usage attribution.
- Establish observability and evaluation practices for non-deterministic systems, including distributed tracing, structured logging, automated evaluations, and cost and latency monitoring.
- Develop and maintain prompt engineering practices, including prompt versioning, regression testing, structured output enforcement, and mitigation of hallucination and prompt-injection risks.
- Diagnose and resolve production incidents across the full stack, and contribute to runbooks, design documentation, and post-incident reviews.
- Collaborate with product, data engineering, and business stakeholders to translate requirements into technical designs, and participate in architecture reviews.
- Follow engineering best practices in code review, automated testing, version control, and CI/CD, and contribute to team technical standards.
- Experience implementing MCP servers or clients, or A2A-based agent interoperability.
- Familiarity with evaluation frameworks for agentic systems, such as LangSmith or Ragas.
- Experience with workflow orchestration platforms such as Prefect, Temporal, Airflow, or Dagster.
- Familiarity with enterprise identity and access management, including Okta, Microsoft Entra ID, SSO, SCIM, and OAuth 2.0.
- Experience with document processing and ingestion at scale, including OCR, parsing of unstructured formats, and multimodal inputs.
- Experience mentoring engineers or leading technical design for a delivery team.
Skills Table: