Job Title
AWS DevOps Engineer / AI Platform Engineer (Hybrid Onsite)
Location
Reading, PA
Duration
6 Months
Interview Type
Virtual / In-Person
Note
Hybrid 2-3 days/wk at client office
Is your candidate willing to relocate for position
Description
Requirements
Experience: 8+ years in Platform Engineering, DevOps, or Site Reliability Engineering (SRE).
Cloud Expertise: Deep proficiency in AWS (IAM, CloudWatch, Bedrock, Lambda).
Observability Tools: Proven experience with Dynatrace, Jaeger, or Honeycomb, and distributed tracing standards.
AI/LLM Interest: Familiarity with the LLM lifecycle, including prompt execution, token usage, and frameworks like LangChain or AgentCore.
Automation: Advanced experience with Terraform and CI/CD pipeline design.
Collaboration: Experience working in an Agile environment with integrated tools like Microsoft Teams and Confluence.
User this when submitting candidates:
Also please check if the next candidate has some experience with at least 50% of below items:
Implementation of Agents on Agentcore runtime
Implementation of Agentic SDLC in Agentcore
Understanding of Strands or any other Agentic AI framework like Langraph, Langchain or Crew AI
Implementation of Bedrock Knowledge Base
Implementation of Knowledge Graph
Implementation of MCP servers in Agentcore
Implementation of Agentcore Gateway
Implementation of Agentcore Identity
Implementation of Agentic AI Observability
Implementation of Agentcore Evaluations
Implementation of AWS Bedrock
Implementation of AWS Bedrock Inference Profile
Implementation of AWS Sagemaker
AWS Services (Cloud) in General
Terraform
Deliverables
Observability
Assess CloudWatch, X-Ray, Bedrock logging, AgentCore traces vs. agentic workflow requirements; produce gap analysis, Setup observability in Dynatrace
Design post-deployment validation pipeline for agents & MCP servers (deployment health + tool registration checks)
Implement distributed tracing & structured logging: LLM decisions, tool selections, sub-agent calls, MCP interactions
Evaluate LangFuse / LiteLLM proxy vs. AWS-native; deliver target-state observability architecture recommendation
Cost Tracking & TCO
Extend tagging taxonomy to cover agent runtimes, MCP servers, vector DBs, Bedrock token consumption per namespace
Design cost visibility model: aggregate agent, MCP, vector DB, and Bedrock token costs per team/department
Build CloudWatch (or equivalent) dashboards for per-team spend; configure AWS Budgets with alerting thresholds
Automate cost reports delivered via email / Microsoft Teams; implement anomaly detection rules
Monitoring & Alerting
Define P1–P4 alerting rules: deployment failures, runtime errors, tool invocation failures, MCP connectivity issues
Integrate alert notifications to Microsoft Teams channels and email; route by resource ownership tags
Author runbooks linked to every alert; publish in Confluence for developer self-service resolution
Evaluate AWS-native vs. third-party monitoring stack; deliver recommendation aligned to observability architecture
Security & Access Control
Assess current IAM + tagging approach for multi-team isolation; identify scalability gaps and risks
Evaluate Cedar policy engine (AgentCore) for fine-grained tool access control; document enterprise-scale gaps
Design scalable ABAC-based identity model for multi-team isolation without IAM policy sprawl; deliver Terraform modules
Top Skills
Skill
Experience required
Platform Engineering / DevOps / SRE
8+ years
AWS Cloud Services
7+ years
Observability Tools
7+ years
AI/LLM Platforms
7+ years
Terraform / CI-CD
7+ years
Agile Collaboration
7+ years
Recruiter
Contact: Sameer - - Seven one nine - two three nine - Five five five five