Responsibilities
- Lead end-to-end delivery of SRE initiatives, including planning, design, execution, testing, deployment, and documentation
- Act as a technical leader for the India-based SRE team by guiding design discussions, collaborating on difficult debugging efforts, and coaching engineers to solve unclear or complex problems independently
- Develop and deploy cloud infrastructure on AWS using tools like Terraform and CloudFormation across services such as EKS, VPC, RDS, IAM, S3, Lambda, CloudFront, SQS, CloudWatch, CloudTrail, and GuardDuty, balancing speed, reliability, and cost efficiency
- Manage Kubernetes environments including cluster maintenance, capacity forecasting, CNI issue resolution, workload scaling, Helm packaging, and GitOps workflows, while creating reusable automation and runbooks
- Improve CI/CD systems using Jenkins with Groovy scripting, GitHub Actions, AWS CodePipeline, ArgoCD, and FluxCD to minimize manual interventions and enhance rollback safety
- Advance observability capabilities by defining infrastructure and architecture guidance that enables teams to use Prometheus, Grafana, and CloudWatch effectively
- Implement FinOps practices such as identifying unused resources, optimizing instance sizing, enforcing tagging policies, and delivering data-backed recommendations for cost savings
- Ensure database reliability for MySQL and PostgreSQL through backup verification, performance optimization, replication monitoring, failover drills, and documented operational procedures
- Enhance security by applying least-privilege access controls, conducting cloud security posture management reviews, monitoring threats via GuardDuty and CloudTrail, managing secrets with Vault and AWS tools, and maintaining audit readiness
- Diagnose and resolve complex production issues involving networking, Kubernetes, compute, databases, and pipelines, then convert solutions into automated responses or documented runbooks
- Produce practical, actionable documentation including architecture decisions, operational procedures, troubleshooting steps, and post-mortem follow-ups that are completed, not just recorded
- Engage daily with U.S.-based SRE leadership on incident reviews, system migrations, roadmap progress, and platform strategy, offering insights and proposals rather than just updates
- Participate in on-call duties and lead post-incident reviews focused on identifying systemic improvements instead of assigning blame