Bangalore, IND; Pune, IND On-site Employment

Zscaler is hiring a Sr. Staff Software Development Engineer - AI Platform

Responsibilities

  • Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints
  • Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases
  • Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR
  • Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)
  • Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews

Requirements

  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3)
  • Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation
  • Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services
  • Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement
  • Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer

Nice to Have

  • Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex
  • Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix
  • Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments
Required Skills
AWS
Got hired remotely?

Get paid like a professional

Remote clients expect company invoices, not personal PayPal requests. Glopay forms an EU partnership that makes you look legitimate while you stay independent.

Professional invoices with EU company details
Compliance handled automatically
Withdraw to any bank account
Income reports for easy tax filing
Create free account
Free signup • 5 min setup
About company
Zscaler
Zscaler (NASDAQ: ZS) operates the world’s largest security cloud, accelerating digital transformation so enterprises can be more agile, efficient, resilient, and secure. The pioneering, AI-powered Zscaler Zero Trust Exchange™ platform protects thousands of enterprise customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location.
All jobs at Zscaler Visit website
Job Details
Department IT Data Strategy
Category other
Posted 12 days ago