Responsibilities
- Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints
- Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases
- Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR
- Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)
- Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews
Requirements
- Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
- 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)
- Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation
- Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services
- Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement
- Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer
Nice to Have
- Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex
- Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix
- Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments